DeepSeek-V4.1-Flash launch cuts every Flash API rate
DeepSeek released V4.1-Flash and cut all six Flash API rates: off-peak cache-hit input fell 57% to $0.003/1M, cache-miss input 32% to $0.15/1M. V4-Pro unchanged.
DeepSeek-V4-Flash (build 0731) on the peak/off-peak card: cache-hit input $0.007/1M off-peak / $0.014/1M peak, cache-miss input $0.22/1M off-peak / $0.44/1M peak, output $0.66/1M off-peak / $1.32/1M peak. A separate experimental model, DeepSeek-V4-Flash-Vision-Exp, was listed on the same rate card. Three model rows on the pricing page.
DeepSeek-V4.1-Flash (model name deepseek-flash): cache-hit input $0.003/1M off-peak / $0.006/1M peak, cache-miss input $0.15/1M off-peak / $0.3/1M peak, output $0.6/1M off-peak / $1.2/1M peak. Vision is a native feature of Flash rather than a separate SKU. Two model rows on the pricing page. DeepSeek-V4-Pro-0813 rates unchanged at $0.022 / $0.66 / $1.98 off-peak, double at peak.
DeepSeek’s Models & Pricing page now lists DeepSeek-V4.1-Flash in place of the previous DeepSeek-V4-Flash-0731 build, and every Flash line item on the peak/off-peak card is cheaper: cache-hit input dropped 57% (from $0.007 to $0.003/1M off-peak), cache-miss input dropped 32% (from $0.22 to $0.15/1M off-peak), and output dropped 9% (from $0.66 to $0.6/1M off-peak). Peak rates remain exactly double the off-peak rates, so each peak figure moved by the same proportion. DeepSeek’s changelog frames it plainly: “With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly.”
The cut partially unwinds the 2026-08-16 time-of-use increase on the Flash side — off-peak input is now back within roughly 7% of the pre-August flat rate, though off-peak output remains a little over 2x it. DeepSeek-V4-Pro’s six rates did not move at all, so the gap between the two models widened: V4-Pro off-peak cache-miss input is now about 4.4x Flash’s, up from 3x.
Two model rows disappeared in the process. The docs footnote states that “the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price.” Vision has been promoted from a separate experimental SKU to a feature row on the main table — supported on V4.1-Flash, “Not supported” on V4-Pro — so image workloads now bill at the standard Flash card. A pinned legacy model name therefore keeps working, but silently changes both the model behind it and the price paid.
DeepSeek released DeepSeek-V4.1-Flash, called as the model name deepseek-flash, and cut every Flash line item on the peak/off-peak card: cache-hit input from $0.007 to $0.003 off-peak ($0.014 to $0.006 peak), cache-miss input from $0.22 to $0.15 off-peak ($0.44 to $0.3 peak), and output from $0.66 to $0.6 off-peak ($1.32 to $1.2 peak). The separate DeepSeek-V4-Flash-0731 and DeepSeek-V4-Flash-Vision-Exp rows were retired into it — vision is now a feature of Flash, and the legacy names still route to V4.1-Flash at the Flash price. V4-Pro's six rates were unchanged, and its announced 2026-09-14 retirement was reversed on 2026-09-11.