Ask
Price change

DeepInfra cuts DeepSeek-V4-Flash-0731 another 25%

DeepInfra pricing

DeepSeek-V4-Flash-0731 fell from $0.08 to $0.06 per 1M input tokens (cached $0.016 to $0.015, output flat at $0.18) — its second tracked cut since its $0.09 catalog debut.

Before

$0.08 in / $0.18 out per 1M tokens ($0.016 cached)

After

$0.06 in / $0.18 out per 1M tokens ($0.015 cached)

DeepSeek-V4-Flash-0731 — the official release build that DeepInfra’s own model card says “supersed[es] the preview version” — was cut again between the 2026-08-28 and 2026-09-07 captures, from $0.08 in / $0.18 out per 1M tokens ($0.016 cached) to $0.06 / $0.18 ($0.015 cached), a 25% reduction on the input rate. This is the model’s second tracked cut since its $0.09 catalog debut: $0.09 (pre-2026-08-04) to $0.08 (2026-08-11) to $0.06 (2026-09-07). The model has appeared in the DeepSeek section of DeepInfra’s headline /pricing page — alongside DeepSeek-V4-Pro, DeepSeek-V4-Flash, DeepSeek-V3.2, and DeepSeek-V3.1 — continuously since at least the 2026-08-28 capture; this cut did not change its listing placement.

From DeepInfra's pricing timeline
Claude catalog restored to seven models; DeepSeek-V4-Flash-0731 cut again

The six Claude models pulled from /pricing on 2026-08-28 return at unchanged rates — claude-opus-5, claude-fable-5 ($10.00/$50.00, again the page's highest per-token rate), claude-sonnet-5, claude-sonnet-4-6, claude-opus-4-7, and claude-opus-4-8 rejoin claude-haiku-4-5. Separately, DeepSeek-V4-Flash-0731 is cut again in the same page's DeepSeek table (input $0.08→$0.06 per 1M, a 25% cut; output flat at $0.18) — its second tracked cut, after the 2026-08-11 move from $0.09 to $0.08 (source: deepinfra.com/pricing, /models, /deepstart 2026-09-07).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart. DeepSeek-V4-Flash-0731 was cut again on 2026-09-07 in the /pricing page's DeepSeek table, from $0.08 to $0.06 per 1M input tokens (output flat at $0.18) — its second tracked cut, after the 2026-08-11 move from $0.09.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Sep 2026
    Claude catalog restored to seven models; DeepSeek-V4-Flash-0731 cut again
  2. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  3. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  4. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  5. Jul 2026
    Flex service tier added at 0.8× base price
Full DeepInfra timeline

More DeepInfra activity

All pricing activity