Ask
Launch

Baseten adds GLM-5.3-Flash and DeepSeek V4 Pro 0813 to Model APIs catalog

Baseten pricing

Baseten grew its Model APIs catalog from 12 to 14 models, adding GLM-5.3-Flash ($0.15/$0.03/$0.50 per 1M tokens) and DeepSeek V4 Pro 0813 ($1.32/$0.132/$3.96); no existing SKU's rate changed.

Before

12 Model APIs SKUs, cheapest input rate $0.10 (GPT OSS 120B) to $0.13 (DeepSeek-V4-Flash-0731); DeepSeek V4 Pro at $1.74 in / $0.145 cache / $3.48 out was the only DeepSeek V4-family SKU.

After

14 Model APIs SKUs adding GLM-5.3-Flash ($0.15 in / $0.03 cache / $0.50 out) and DeepSeek V4 Pro 0813 ($1.32 in / $0.132 cache / $3.96 out, a lower-input/higher-output dated variant of DeepSeek V4 Pro); all 12 prior SKUs unchanged.

Proof of change

Baseten's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Aug 4, 2026 Captured Aug 28, 2026 · 24 days apart
main-minute
$0.15 $0.03 $1.32 $3.96

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Aug 4, 2026
Captured Aug 28, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 2 other pages we compared
main-hour
$0.15 $1.32 $3.96

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Aug 4, 2026
Captured Aug 28, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

docs
Captured Aug 4, 2026
Captured Aug 28, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Three weeks after renaming DeepSeek V4 to DeepSeek V4 Pro alongside the Flash and Inkling-Small additions, Baseten grew its Model APIs catalog again — from twelve models to fourteen. GLM-5.3-Flash launched at $0.15 input / $0.03 cache input / $0.50 output per 1M tokens, promoted by a new homepage banner (“Try the new GLM-5.3 Flash today. Frontier intelligence at a fraction of the cost.”) that replaced the prior “Try the new DeepSeek V4 Flash today” promo — the second consecutive banner cycle built around a new cheap-tier Model API launch rather than a funding milestone.

DeepSeek V4 Pro 0813 joined alongside it as a second, dated DeepSeek V4 variant: its $1.32 input rate is 24% below the original DeepSeek V4 Pro’s $1.74, but its $3.96 output rate is 14% above DeepSeek V4 Pro’s $3.48 — the first Model APIs addition on this rate card to cut one rate while raising another on the same SKU family. No rate on any of the twelve prior SKUs moved, and the dedicated GPU, CPU, Training rate cards and the Basic/Pro/Enterprise tier structure were all unchanged.

From Baseten's pricing timeline
Model APIs Catalog Grows to Fourteen Models

Three weeks after the DeepSeek V4 Pro rename, Baseten added GLM-5.3-Flash ($0.15 input / $0.03 cache input / $0.50 output per 1M tokens) and DeepSeek V4 Pro 0813 ($1.32 / $0.132 / $3.96 — a dated DeepSeek V4 Pro variant with a 24% lower input rate but 14% higher output rate). No rate on any of the twelve prior SKUs changed, and the homepage banner switched from promoting DeepSeek V4 Flash to promoting GLM-5.3-Flash. Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity