DeepInfra cuts DeepSeek-V4-Flash-0731 pricing about 11%
DeepInfra cut DeepSeek-V4-Flash-0731 pricing ~11%: $0.09/$0.18 to $0.08/$0.18 per 1M tokens (in/out), a second catalog-only model repriced outside the main /pricing page.
$0.09 in / $0.18 out per 1M tokens ($0.018 cached)
$0.08 in / $0.18 out per 1M tokens ($0.016 cached)
Proof of change
DeepInfra's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 2 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
- These prices come from our capture alone — they were not confirmed against an independent second source.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
DeepSeek-V4-Flash-0731 — the official release build of DeepSeek-V4-Flash, “superseding the preview version” per DeepInfra’s own model description — dropped from $0.09 in / $0.18 out per 1M tokens ($0.018 cached) to $0.08 in / $0.18 out ($0.016 cached) between the 2026-08-04 and 2026-08-11 captures of DeepInfra’s /models and /deepstart catalogs, an ~11% cut on the input and cached rates with the output rate held flat. A neighboring model, Kimi-K2.7-Code, saw a smaller simultaneous cut ($0.74/$3.50/$0.15 cached to $0.68/$3.40/$0.136 cached per 1M), while every other featured catalog model checked in the same comparison held its rate exactly. Like the GLM-5.2 cut logged on 2026-08-04, DeepSeek-V4-Flash-0731’s price is not shown on DeepInfra’s main /pricing page — it surfaces only via the /models Featured carousel and the /deepstart startup-credits page, reinforcing that a complete read of DeepInfra’s rate card now requires checking more than the headline pricing page.
GLM-5.2 (Z-AI's flagship long-horizon agentic model, 1M-token context) is cut roughly 20% across all three published rates — $0.93 in / $3.00 out / $0.18 cached to $0.75 / $2.40 / $0.14 per 1M tokens — while every other Featured model checked in the same comparison (DeepSeek-V4-Pro, DeepSeek-V4-Flash, Kimi-K2.6/K2.7-Code, Nemotron-3-Ultra, Qwen3-Max, Qwen3.6-35B-A3B, GLM-5.1, MiMo-V2.5-Pro) holds its rate exactly. GLM-5.2 is not shown on the /pricing page at all — it surfaces only in the /models Featured carousel — making this the first tracked price move that lives entirely outside DeepInfra's main pricing surface (source: deepinfra.com/models 2026-08-04).