Ask
Price change

DeepSeek-V4.1-Flash launch cuts every Flash API rate

DeepSeek pricing

DeepSeek released V4.1-Flash and cut all six Flash API rates: off-peak cache-hit input fell 57% to $0.003/1M, cache-miss input 32% to $0.15/1M. V4-Pro unchanged.

Before

DeepSeek-V4-Flash (build 0731) on the peak/off-peak card: cache-hit input $0.007/1M off-peak / $0.014/1M peak, cache-miss input $0.22/1M off-peak / $0.44/1M peak, output $0.66/1M off-peak / $1.32/1M peak. A separate experimental model, DeepSeek-V4-Flash-Vision-Exp, was listed on the same rate card. Three model rows on the pricing page.

After

DeepSeek-V4.1-Flash (model name deepseek-flash): cache-hit input $0.003/1M off-peak / $0.006/1M peak, cache-miss input $0.15/1M off-peak / $0.3/1M peak, output $0.6/1M off-peak / $1.2/1M peak. Vision is a native feature of Flash rather than a separate SKU. Two model rows on the pricing page. DeepSeek-V4-Pro-0813 rates unchanged at $0.022 / $0.66 / $1.98 off-peak, double at peak.

DeepSeek’s Models & Pricing page now lists DeepSeek-V4.1-Flash in place of the previous DeepSeek-V4-Flash-0731 build, and every Flash line item on the peak/off-peak card is cheaper: cache-hit input dropped 57% (from $0.007 to $0.003/1M off-peak), cache-miss input dropped 32% (from $0.22 to $0.15/1M off-peak), and output dropped 9% (from $0.66 to $0.6/1M off-peak). Peak rates remain exactly double the off-peak rates, so each peak figure moved by the same proportion. DeepSeek’s changelog frames it plainly: “With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly.”

The cut partially unwinds the 2026-08-16 time-of-use increase on the Flash side — off-peak input is now back within roughly 7% of the pre-August flat rate, though off-peak output remains a little over 2x it. DeepSeek-V4-Pro’s six rates did not move at all, so the gap between the two models widened: V4-Pro off-peak cache-miss input is now about 4.4x Flash’s, up from 3x.

Two model rows disappeared in the process. The docs footnote states that “the legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted, but the corresponding models have been retired, their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price.” Vision has been promoted from a separate experimental SKU to a feature row on the main table — supported on V4.1-Flash, “Not supported” on V4-Pro — so image workloads now bill at the standard Flash card. A pinned legacy model name therefore keeps working, but silently changes both the model behind it and the price paid.

From DeepSeek's pricing timeline
DeepSeek-V4.1-Flash Launches — Flash Rates Cut Across the Board

DeepSeek released DeepSeek-V4.1-Flash, called as the model name deepseek-flash, and cut every Flash line item on the peak/off-peak card: cache-hit input from $0.007 to $0.003 off-peak ($0.014 to $0.006 peak), cache-miss input from $0.22 to $0.15 off-peak ($0.44 to $0.3 peak), and output from $0.66 to $0.6 off-peak ($1.32 to $1.2 peak). The separate DeepSeek-V4-Flash-0731 and DeepSeek-V4-Flash-Vision-Exp rows were retired into it — vision is now a feature of Flash, and the legacy names still route to V4.1-Flash at the Flash price. V4-Pro's six rates were unchanged, and its announced 2026-09-14 retirement was reversed on 2026-09-11.

About DeepSeek
api-docs.deepseek.com ↗

DeepSeek offers a free web chat product and a pay-per-token API, billed since 2026-08-16 on a peak/off-peak rate card: DeepSeek-V4.1-Flash (general-purpose, from $0.003/1M cache-hit input off-peak / $0.006 peak, $0.15 cache-miss in off-peak / $0.3 peak, $0.6 out off-peak / $1.2 peak) and DeepSeek-V4-Pro (higher-capacity, from $0.022/1M cache-hit input off-peak / $0.044 peak, $0.66 cache-miss in off-peak / $1.32 peak, $1.98 out off-peak / $3.96 peak) — both with a 1M-token context window and dramatically cheaper than equivalent OpenAI and Anthropic models.

Pricing model freemiumpure usage
Billing units tokensapi calls
Sales motion self serveplg
Free tier
Yes
Commits
None
Transparency
public

DeepSeek pricing history

  1. Sep 2026
    DeepSeek-V4.1-Flash Launches — Flash Rates Cut Across the Board
  2. Aug 2026
    DeepSeek Announces Peak/Off-Peak Pricing — Effective Aug 16, 2026
  3. Mar 2025
    DeepSeek-V3-0324 Update — Improved Coding
  4. Jan 2025
    Nvidia Stock Drops 17% Following R1 Release
  5. Jan 2025
    DeepSeek-R1 Released — Reasoning Model, MIT Open-Source
Full DeepSeek timeline

More DeepSeek activity

All pricing activity