DeepInfra hikes its cheapest small model output price 33%
DeepInfra raised Llama-3.1-8B-Instruct-Turbo output price 33% to $0.04 per 1M tokens (from $0.03) and cut gemma-4-31B-it-turbo to $0.09/$0.34 per 1M.
Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.03 out per 1M tokens; gemma-4-31B-it-turbo $0.12 in / $0.37 out per 1M tokens.
Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.04 out per 1M tokens (plus 33% on output); gemma-4-31B-it-turbo $0.09 in / $0.34 out per 1M tokens (minus 25% input, minus 8% output).
DeepInfra’s price list moved on two individual model SKUs between the 2026-07-21 and 2026-07-29 captures. Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra’s own catalog repeatedly cite as the cheapest small-model reference — rose from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens, a 33 percent increase on the output rate. In the same capture, gemma-4-31B-it-turbo was cut from $0.12/$0.37 to $0.09/$0.34 per 1M tokens. Headline per-GPU-hour rates, DeepCluster reserved pricing, usage tiers, and the three-step Flex/Standard/Priority service tier were all unchanged. This is the second rate increase DeepInfra has published since its 2026-07-14 GPU-hour hike broke an eighteen-month streak of price cuts, suggesting the July reversal was not a one-off.
DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra's own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update. GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged; this is the second rate increase since the 2026-07-14 reversal, suggesting it was not a one-off (source: deepinfra.com/pricing 2026-07-29).