Ask
Price change

DeepInfra hikes its cheapest small model output price 33%

DeepInfra pricing

DeepInfra raised Llama-3.1-8B-Instruct-Turbo output price 33% to $0.04 per 1M tokens (from $0.03) and cut gemma-4-31B-it-turbo to $0.09/$0.34 per 1M.

Before

Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.03 out per 1M tokens; gemma-4-31B-it-turbo $0.12 in / $0.37 out per 1M tokens.

After

Meta-Llama-3.1-8B-Instruct-Turbo $0.02 in / $0.04 out per 1M tokens (plus 33% on output); gemma-4-31B-it-turbo $0.09 in / $0.34 out per 1M tokens (minus 25% input, minus 8% output).

Proof of change

DeepInfra's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 21, 2026 Captured Jul 29, 2026 · 8 days apart
startup-credits
$0.08 $0.40 $2.00 $0.13 $0.38
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 3 other pages we compared
models
$2.00 $0.00
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

main
$0.37
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

enterprise
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

DeepInfra’s price list moved on two individual model SKUs between the 2026-07-21 and 2026-07-29 captures. Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra’s own catalog repeatedly cite as the cheapest small-model reference — rose from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens, a 33 percent increase on the output rate. In the same capture, gemma-4-31B-it-turbo was cut from $0.12/$0.37 to $0.09/$0.34 per 1M tokens. Headline per-GPU-hour rates, DeepCluster reserved pricing, usage tiers, and the three-step Flex/Standard/Priority service tier were all unchanged. This is the second rate increase DeepInfra has published since its 2026-07-14 GPU-hour hike broke an eighteen-month streak of price cuts, suggesting the July reversal was not a one-off.

From DeepInfra's pricing timeline
Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%

DeepInfra raises Meta-Llama-3.1-8B-Instruct-Turbo — the model this page and DeepInfra's own catalog cite as the cheapest small-model reference — from $0.02 in / $0.03 out to $0.02 in / $0.04 out per 1M tokens (+33% on output), while cutting gemma-4-31B-it-turbo from $0.12/$0.37 to $0.09/$0.34 per 1M in the same update. GPU-hour, DeepCluster, usage-tier, and service-tier rates are unchanged; this is the second rate increase since the 2026-07-14 reversal, suggesting it was not a one-off (source: deepinfra.com/pricing 2026-07-29).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart. DeepSeek-V4-Flash-0731 was cut again on 2026-09-07 in the /pricing page's DeepSeek table, from $0.08 to $0.06 per 1M input tokens (output flat at $0.18) — its second tracked cut, after the 2026-08-11 move from $0.09.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Sep 2026
    Claude catalog restored to seven models; DeepSeek-V4-Flash-0731 cut again
  2. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  3. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  4. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  5. Jul 2026
    Flex service tier added at 0.8× base price
Full DeepInfra timeline

More DeepInfra activity

All pricing activity