Ask
Packaging

Baseten trims Model APIs catalog to eight models and adds Inkling

Baseten pricing

Baseten cut its Model APIs rate card from 11 models to 8, retiring GLM 5.1, GLM 5, Kimi K2.5 and Nemotron 3 Super, and added Inkling at $1.00/$4.05 per 1M tokens.

Before

11 Model APIs SKUs: GLM 5.2, GLM 5.1, GLM 5, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, Kimi K2.5, NVIDIA Nemotron 3 Ultra, NVIDIA Nemotron 3 Super, DeepSeek V4, GPT OSS 120B.

After

8 Model APIs SKUs: Inkling ($1.00 in / $0.17 cache / $4.05 out), GLM 5.2, GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.

Proof of change

Baseten's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 6, 2026 Captured Jul 21, 2026 · 15 days apart
main-hour
$1.30 $4.30 $3.15 $3.00 $0.30 $0.75 $1.00 $0.17 $4.05
Captured Jul 6, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 2 other pages we compared
main-minute
$1.30 $4.30 $0.20 $3.15 $3.00 $0.30 $1.00 $0.17 $4.05
Captured Jul 6, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

docs
Captured Jul 6, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Between the 2026-07-06 and 2026-07-21 captures of baseten.co/pricing, four models disappeared from the Model APIs rate card in a single sweep — GLM 5.1 ($1.30 in / $4.30 out), GLM 5 ($0.95 / $3.15), Kimi K2.5 ($0.60 / $3.00) and NVIDIA Nemotron 3 Super ($0.30 / $0.75) — while Inkling was added at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens.

The retirements pull the cheap end of the catalog upward: with Nemotron 3 Super gone, the lowest published input rate outside GPT OSS 120B ($0.10) is now GLM 4.7 and Nemotron 3 Ultra at $0.60, and the cheapest published cache-input rate rises from $0.06 to $0.12. Baseten documents a deprecation policy for Model APIs, so the sweep reads as scheduled catalog pruning rather than a repricing.

No other pricing surface moved. Dedicated deployment GPU rates (T4 $0.01052/min through B200 $0.16633/min), CPU instance rates, the on-demand Training rate card, and the Basic / Pro / Enterprise tier structure are all unchanged from the prior capture.

From Baseten's pricing timeline
Model APIs Catalog Pruned to Eight Models; Inkling Added

Baseten cut its Model APIs rate card from eleven SKUs to eight, retiring GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super, and added Inkling at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens. No surviving model's rate changed, but the retirements removed the cheap end of the catalog: the lowest published cache-input rate doubled from $0.06 to $0.12. Dedicated GPU, CPU, Training and Basic/Pro/Enterprise surfaces were unchanged.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity