Ask
Launch

Baseten launches DeepSeek V4 Flash and Inkling-Small, renames DeepSeek V4 to Pro

Baseten pricing

Baseten added DeepSeek V4 Flash and Inkling-Small to its Model APIs catalog (now 12 models), renaming DeepSeek V4 to DeepSeek V4 Pro at unchanged rates.

Before

10 Model APIs SKUs: Kimi K3, GLM-5.2 Fast, Inkling, GLM-5.2, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4 ($1.74 in / $0.145 cache / $3.48 out), GPT OSS 120B.

After

12 Model APIs SKUs adding DeepSeek-V4-Flash-0731 ($0.13 in / $0.028 cache / $0.26 out) and Inkling-Small ($0.50 / $0.10 / $1.20); DeepSeek V4 renamed DeepSeek V4 Pro (rates unchanged: $1.74 / $0.145 / $3.48).

Eight days after growing its Model APIs catalog to ten models with Kimi K3 and GLM-5.2 Fast, Baseten expanded the roster again to twelve — this time at the cheap end of the rate card. DeepSeek-V4-Flash-0731 launched as the platform’s lowest-priced frontier-derived model at $0.13 input / $0.028 cache input / $0.26 output per 1M tokens, promoted by a homepage banner (“Try the new DeepSeek V4 Flash today. Frontier intelligence at a fraction of the cost.”) that replaced the prior “Announcing our Series F” funding banner. Inkling-Small joined alongside it at $0.50 / $0.10 / $1.20, a smaller and cheaper sibling to the existing Inkling line.

The existing DeepSeek V4 SKU was simultaneously renamed DeepSeek V4 Pro to disambiguate it from the new Flash variant — its $1.74 input, $0.145 cache input, and $3.48 output rates did not change. No other pricing surface moved: dedicated deployment GPU/CPU rates, the on-demand Training rate card, and the Basic/Pro/Enterprise tier structure are all unchanged from the prior capture.

From Baseten's pricing timeline
Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro

Six days after growing to ten models, Baseten added DeepSeek-V4-Flash-0731 ($0.13 input / $0.028 cache input / $0.26 output per 1M tokens — the platform's cheapest model to date) and Inkling-Small ($0.50 / $0.10 / $1.20), and renamed the existing DeepSeek V4 SKU to DeepSeek V4 Pro at unchanged rates ($1.74 / $0.145 / $3.48). The homepage banner switched from an 'Announcing our Series F' funding promo to a DeepSeek V4 Flash product promo. Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity