Ask
Packaging

DeepSeek adds Responses API to V4-Flash, announces future peak-hour surcharge

DeepSeek pricing

DeepSeek's pricing docs now list a Responses API feature (Flash-only today) and a footnoted plan to charge 2x during Beijing-time peak hours, though no effective date or per-token price has changed yet.

Before

Flat per-token rate card with no time-of-use pricing; no Responses API row on the model table.

After

New Responses API row (deepseek-v4-flash supported now, deepseek-v4-pro coming ~early Aug 2026); a new footnote discloses an upcoming peak/off-peak policy charging 2x the regular rate during 9:00-12:00 and 14:00-18:00 Beijing Time, effective date not yet announced.

DeepSeek’s official Models & Pricing page (api-docs.deepseek.com) picked up two structural additions between the 2026-07-29 and 2026-08-05 captures, while every per-token rate stayed unchanged ($0.14/$0.0028/$0.28 for V4-Flash; $0.435/$0.003625/$0.87 for V4-Pro, cache-miss/cache-hit/output per 1M tokens).

First, a Responses API feature row now appears on the model comparison table, live today for deepseek-v4-flash only — the docs note deepseek-v4-pro support is coming “in early August 2026.” Second, and more consequential, DeepSeek disclosed a forthcoming time-of-use pricing policy: during peak hours (9:00-12:00 and 14:00-18:00 Beijing Time daily), “prices will be 2x the regular prices, applicable to all billing items.” No effective date has been set — DeepSeek says it is “subject to the official announcement” — so today’s flat rate card still applies. This reverses the direction of DeepSeek’s now-discontinued 2025-era off-peak discount into a peak-hour surcharge, a notable shift for a company whose pricing has so far moved in only one direction: down.

From DeepSeek's pricing timeline
DeepSeek Announces Peak/Off-Peak Pricing — Effective Aug 16, 2026

DeepSeek replaced its vague "significant price increase" footnote with a concrete time-of-use rate card: peak-hour rates (01:00-04:00 and 06:00-10:00 UTC) roughly 3-11x current per-token prices depending on the line item, with off-peak rates at half of peak, effective 16:00 UTC on 2026-08-16. V4-Flash cache-miss input moves from a flat $0.14 to $0.22 off-peak / $0.44 peak; V4-Pro cache-miss input moves from a flat $0.435 to $0.66 off-peak / $1.32 peak. DeepSeek-V4-Pro also picked up a dated build suffix (-0813) and gained Responses API support, reaching feature parity with V4-Flash.

About DeepSeek
api-docs.deepseek.com ↗

DeepSeek offers a free web chat product and a pay-per-token API, billed since 2026-08-16 on a confirmed peak/off-peak rate card: DeepSeek-V4-Flash (general-purpose, from $0.007/1M cache-hit input off-peak / $0.014 peak, $0.22 cache-miss in off-peak / $0.44 peak, $0.66 out off-peak / $1.32 peak), DeepSeek-V4-Pro (higher-capacity, from $0.022/1M cache-hit input off-peak / $0.044 peak), and the experimental DeepSeek-V4-Flash-Vision-Exp (billed on V4-Flash's rate card) — all with a 1M-token context window and dramatically cheaper than equivalent OpenAI and Anthropic models.

Pricing model freemiumpure usage
Billing units tokensapi calls
Sales motion self serveplg
Free tier
Yes
Commits
None
Transparency
public

DeepSeek pricing history

  1. Aug 2026
    DeepSeek Announces Peak/Off-Peak Pricing — Effective Aug 16, 2026
  2. Mar 2025
    DeepSeek-V3-0324 Update — Improved Coding
  3. Jan 2025
    Nvidia Stock Drops 17% Following R1 Release
  4. Jan 2025
    DeepSeek-R1 Released — Reasoning Model, MIT Open-Source
  5. Dec 2024
    DeepSeek-V3 Released — Frontier Performance at $0.27/1M
Full DeepSeek timeline

More DeepSeek activity

All pricing activity