Ask
Price change

Cached-input pricing comes to Model APIs

Baseten pricing

Baseten adds discounted cached-input pricing to its multi-tenant Model APIs.

From Baseten's pricing timeline
Cached Input Pricing on Model APIs

Baseten added a dedicated Cache Input column to the Model APIs rate card — prompts re-using prefixes from prior requests within a session window are billed far below the standard input rate. This narrows the gap against first-party caching offerings from OpenAI and Anthropic.

About Baseten
baseten.co ↗

Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.

Free tier
Yes
Commits
Available
Transparency
public

Baseten pricing history

  1. Aug 2026
    Model APIs Catalog Grows to Fourteen Models
  2. Aug 2026
    Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro
  3. Jul 2026
    Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%
  4. Jul 2026
    Model APIs Catalog Pruned to Eight Models; Inkling Added
  5. Feb 2026
    Cached Input Pricing on Model APIs
Full Baseten timeline

More Baseten activity

All pricing activity