DeepInfra hikes its cheapest small model output price 33%
DeepInfra raised Llama-3.1-8B-Instruct-Turbo output price 33% to $0.04 per 1M tokens (from $0.03) and cut gemma-4-31B-it-turbo to $0.09/$0.34 per 1M.
Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.03 out per 1M tokens; gemma-4-31B-it-turbo $0.12 in / $0.37 out per 1M tokens.
Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.04 out per 1M tokens (plus 33% on output); gemma-4-31B-it-turbo $0.09 in / $0.34 out per 1M tokens (minus 25% input, minus 8% output).
Proof of change
DeepInfra's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 3 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
DeepInfra’s price list moved on two individual model SKUs between the 2026-07-21 and 2026-07-29 captures. Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra’s own catalog repeatedly cite as the cheapest small-model reference — rose from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens, a 33 percent increase on the output rate. In the same capture, gemma-4-31B-it-turbo was cut from $0.12/$0.37 to $0.09/$0.34 per 1M tokens. Headline per-GPU-hour rates, DeepCluster reserved pricing, usage tiers, and the three-step Flex/Standard/Priority service tier were all unchanged. This is the second rate increase DeepInfra has published since its 2026-07-14 GPU-hour hike broke an eighteen-month streak of price cuts, suggesting the July reversal was not a one-off.
DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra's own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update. GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged; this is the second rate increase since the 2026-07-14 reversal, suggesting it was not a one-off (source: deepinfra.com/pricing 2026-07-29).