Ask
Sharpens 21 companies · First observed July 2024 · Updated September 2026 Explore in the graph

Prompt Caching as a Structural Pricing Feature

Quick answer

Nine AI API vendors now publish a discounted cached-input rate — 50-80% below the standard input price — as a structural pricing tier. Caching rewards workloads with stable, repeated system prompts and raises switching costs for RAG and agent-heavy applications.

9 vendors publish discounted cached-input pricing

What's happening — and why

What's happening: prompt and context caching has become a standard pricing feature across frontier LLM APIs. Nine corpus companies publish a distinct cached-input price that applies when input tokens match a previously stored prefix. Discounts range from 50% (OpenAI) to 75-80% (Anthropic, Google, DeepSeek).

Why: caching is an efficiency win for both sides. For vendors, a cached prompt avoids re-encoding the same context — lower compute cost. For buyers, applications that reuse a stable system prompt, RAG document set, or codebase context can cut input costs dramatically. The vendor passes the savings while keeping per-token rates for fresh, diverse inputs.

The strategic implication is stickiness. A workload heavily invested in a vendor's caching structure — with optimal cache-key design and stored contexts — is expensive to migrate. The cached-input discount is both a cost reduction and a switching-cost mechanism.

How it works

Uncached input Full price e.g. $15/1M tokens Cached input (same prefix repeated) Discounted $3.75/1M (75% off) Anthropic OpenAI 50% Google 75% DeepSeek 74% → Groq, Fireworks → Together, Baseten Minimum cache size: typically 1k–32k tokens
Nine vendors offer 50-80% off for cached input tokens — standard at frontier labs, spreading to inference platforms.

Evidence over time

30 supporting · 5 counter — hover or tap a point for detail, click to jump to the row.

supports ↑ challenges ↓ 2024 2025 2026
supporting evidence counterexample

Evidence

Company Date What happened
anthropic Aug 2024 Prompt Caching launched August 2024: $3.75/1M cached input (vs $15/1M full) for Claude 3.5 Sonnet — 75% discount on repeated context
openai Oct 2024 Context caching launched October 2024: 50% discount on cached input tokens across GPT-4o and GPT-4o-mini
google-gemini Jul 2024 Context caching in Gemini API: $0.01875/1M cached input (vs $0.075) for Gemini 1.5 Flash — 75% discount; minimum 32k token cache
deepseek Jan 2025 DeepSeek V3 cache hits priced at $0.07/1M (vs $0.27 input) — 74% off for cache hits
together-ai Jun 2025 Prompt caching available across Llama and Qwen models; discount rates on repeat context
groq Sep 2025 On-demand context caching for qualifying models; cache storage free for limited windows
fireworks-ai Sep 2025 Context caching supported on Llama 3.x family; cache hits discounted vs fresh input
baseten Feb 2026 Cached-input pricing added to Model APIs — discounted rate for repeated prefill context in multi-tenant inference
replicate Jun 2025 Prediction warmup and model caching features reduce cold-start costs; effectively a caching discount for warm models
deepinfra Jun 2026 Shows a discounted cached-input price inline next to the standard rate on many per-token rows across hundreds of open models — cache economics on the public price list rather than behind docs. Also the clearest cost-forecasting problem the layer creates: every model row now carries three rates (input, output, cached input).
xai Jul 2026 Cut grok-4.5 cached input from $0.50 to $0.30 per 1M — a 40% cut with uncached rates held at $2.00 input / $6.00 output (500k context). Deepens the cache discount from 4× to 6.7×, closing on grok-4.3's $0.20-against-$1.25. Also publishes a separate long-context cached rate at $0.60/1M. A repricing that touched only the cached line.
moonshot-ai Jul 2026 Kimi K3 lands as flagship at $3.00/1M cache-miss input, $0.30/1M cache-hit, $15.00/1M output on a 1,048,576-token context — a clean 10× cache discount at the top of the card. Kimi K2.7 Code at $0.95/$0.19 cache hit (5×), and its HighSpeed variant doubles EVERY rate including the cache-hit line ($0.38) — throughput is priced on the cached tokens too.
minimax Jul 2026 MiniMax-M2.7 and the frontier M3 (MSA sparse attention, 1M context) both bill $0.30/M input and $1.20/M output with cache reads at $0.06/M — a 5× discount published as a standard column on the international USD API card.
sambanova Jul 2026 NEW ADOPTER. Added a first-ever 'Cached Input Tokens' column to SambaCloud's public rate card, with MiniMax-M2.7 at $0.06/1M cached vs $0.60/1M uncached — a 10× discount. Notably only 1 of the 6 models on the trimmed card carries a cached rate, so caching arrived as a per-model capability rather than a platform-wide policy.
novita-ai Jul 2026 Reseller pass-through, verified: cache-read rates printed inline across a 196-model catalog — Kimi K3 at $0.3/M cache read against $3/M input, Tencent Hy3 at $0.035/M cache read against $0.14/M input (4×). Confirms the layer propagating down to aggregators that resell upstream capacity.
zhipu-ai Jul 2026 Publishes cached input as low as $0.01 per 1M tokens, with $0.11 on the GLM-4.x flagships and $0.26 on GLM-5.2 — and separately lists a 'Cached Input Storage' line marked 'Limited-time Free' on every model that supports it. The first corpus vendor to name cache storage as its own billable resource, currently zero-rated.
anthropic Jul 2026 Claude Opus 5 launches at $5/$25 per MTok with the cache multipliers carried over unchanged: $0.50/MTok read (10% of base input, a 10× discount) and $6.25/MTok for a 5-minute cache WRITE (1.25× base input); a 1-hour cache write costs 2× base. Sonnet 5's read is $0.20/1M at introductory pricing. Anthropic's own worked example puts caching at a 66% cost reduction on a long-context workload and 83% when combined with the 50%-off Batch API.
fireworks-ai Jul 2026 Split serverless into three named serving paths — Standard, Priority (via service_tier) and Fast (separate model ID, 100+ tok/s) — each with its own published input / cached-input / output rate. Cached input is now a mandatory column replicated across every quality-of-service tier, not a single per-model rate.
braintrust Jul 2026 The layer reaches the application tier: Braintrust — an eval/observability platform, not an inference provider — publishes its own cached-input rate as part of the renamed 'Model credits' pool, with GLM-5.2 overage at $1.40/mtok input, $0.26/mtok CACHED input and $4.40/mtok output. A 5.4× cache discount printed on a tooling rate card.
poolside Jul 2026 Cache hit rate has become a marketing claim, not just a price line: Poolside advertises effective spend well under its $0.10/1M list on Laguna S 2.1 'driven by a 95–97% cache-hit rate on repeated context.' The vendor is now selling the cache discount as the headline number rather than the rate card.
baseten Jul 2026 The cleanest cached-line-only move in the corpus, and it lands exactly on 10×. GLM-5.2's cache-input rate was cut 46%, from $0.26 to $0.14 per 1M, with input and output both held at $1.40 and $4.40 — deepening the discount from 5.4× to 10.0×. New flagship Kimi K3 arrived in the same capture at $3.00 in / $0.30 cache / $15.00 out (10× again), and GLM-5.2 Fast at $2.10 / $0.21 / $6.60 (also 10×). Three SKUs on one card at precisely one-tenth.
deepseek Aug 2026 REVERSAL, and a new multiplier. DeepSeek confirmed peak/off-peak API pricing effective 16:00 UTC on 2026-08-16 (peak 01:00-04:00 and 06:00-10:00 UTC, off-peak exactly half), three days after warning of an unspecified 'significant' increase. Cache-hit input — DeepSeek's cheapest and most aggressively marketed rate, and the line it used as its price-CUT vehicle in April 2026 — took the biggest relative increase of any item: V4-Flash from a flat $0.0028/1M to $0.007 off-peak (2.5×) and $0.014 peak (5×); V4-Pro from $0.003625/1M to $0.022 off-peak and $0.044 peak, over 12×. The discount ratios collapse from 50× to 31× on V4-Flash and from 120× to 30× on V4-Pro. Two consequences: the corpus's deepest cache discounts converged downward toward the 10-30× pack, and the cached line now inherits a TIME-OF-DAY multiplier alongside serving path and context band. No grace period or legacy-rate opt-out was published.
weights-biases Aug 2026 The cached line moving AGAINST its own headline, in one revision. W&B cut DeepSeek V4-Pro Serverless Inference input 34% ($1.74 → $1.15 per 1M) and output 27% ($3.48 → $2.55) while RAISING cached input 43% ($0.14 → $0.20). The discount shrank from 12.4× to 5.75× on a model that got materially cheaper to use uncached. The prior framing had vendors cutting the cached line while holding headlines; this is the mirror case, and it means a headline price cut can quietly raise the bill for a high-hit-rate workload.
perplexity-ai Aug 2026 NEW ADOPTER, and it publishes a cached rate that saves nothing on one SKU. Perplexity's Gateway API — its first developer surface priced with real hosting margin rather than at-cost third-party resale — launched with a cache-read column on every model: perplexity/kimi-k3 at $3.00 in / $15.00 out / $0.30 cache-read (10×), perplexity/glm-5.2 at $1.40 / $4.40 / $0.14 (10×), perplexity/deepseek-v4-flash-0731 at $0.13 / $0.26 / $0.028 (4.6×). By 2026-08-14 two NVIDIA models joined: nemotron-3.5-lightning-30b-a3b at $0.0115 / $0.17 / $0.00115 (10×) and nemotron-3-ultra-550b-a55b at $0.25 / $2.50 / $0.25 — a cache-read rate identical to input, a 1× 'discount'. The column is now mandatory even when the number behind it grants nothing.
cursor Aug 2026 Shows the pass-through tier CANNOT use the cache as a lever. Cursor bills its Other Models pool at the model's own API rate, so when GPT-5.6 Luna was repriced the cached lines moved in exact lockstep with the headline: input $1 → $0.20, cache write $1.25 → $0.25, cache read $0.10 → $0.02, output $6 → $1.20 — every dimension down ~80%, the discount ratio unchanged at 10×. Augment Code (2026-08-11), Vercel/v0 (2026-08-11) and Glean (2026-08-04) published the identical proportional move on the same model in the same week. Independent cached-line pricing is a privilege of vendors who set their own rates.
fireworks-ai Aug 2026 Cached rates now ship with every new SKU as a matter of course. Two models were added with full three-rate cards: NVIDIA Nemotron 3.5 Lightning 30B A3B at $0.05 input / $0.01 cached / $0.20 output per 1M (5×) — the cheapest published input rate on the entire card, undercutting GPT OSS 20B at $0.07 — and Muse Glimmer 30B at $0.35 / $0.04 / $1.50 Standard (8.75×) with a Priority path at $0.525 / $0.06 / $2.25. The cached line is replicated per serving path, not per model.
poolside Aug 2026 The marketing claim got a measured number. OpenRouter's trailing 7-day weighted average now puts Poolside's effective post-cache spend at $0.011 input / $0.179 output per 1M for Laguna S 2.1 against a $0.09/$0.18 list — a 97.7% observed cache-hit rate, above the 95-97% Poolside advertises. Laguna XS 2.1 runs $0.036/$0.119 against $0.06/$0.12. The list price is now roughly 8× the price the median caller actually pays on input.
Cursor Aug 2026 Reconfirms the at-cost-reseller sub-mechanic on a second upstream and a second SKU family. Cursor cut Claude Sonnet 5 from $3 / $3.75 cache-write / $0.3 cache-read / $15 to $2 / $2.5 / $0.2 / $10 — every line including both cache rates moving in exact lockstep with Anthropic's introductory card — and did the same for GPT-5.6 Sol ($5/$6.25/$0.5/$30 to $4/$5/$0.4/$20). A pass-through reseller cannot run an independent cached line, so its cache ratio is fixed at the upstream's.
Baseten Aug 2026 Two new SKUs land with cache rates published from day one, at very different ratios: GLM-5.3-Flash at $0.15 input / $0.03 cached / $0.50 output (a 5x cache discount) and DeepSeek V4 Pro 0813 at $1.32 / $0.132 / $3.96 (a clean 10x). Catalog went 12 to 14 Model APIs SKUs with all 12 prior rates unchanged. The 10x frontier this trend identifies is now the default for new dated variants, while a cheap flash model ships at half that ratio.
DeepInfra Aug 2026 Cached rates discounted in lockstep with a promotional headline cut, preserving the ratio rather than moving it. MiMo-V2.5-Pro launched at 61% off — $1.00/$3.00 list with $0.20 cached struck to $0.39/$1.17 with $0.078 cached — holding a 5x cache ratio through the promo. GLM-5.2 took a 35% promo tag on 2026-08-28, moving $0.75/$2.40 with $0.14 cached to $0.488/$1.56 with $0.091 cached. So a promo can propagate to the cached line without changing the discount depth, which is the opposite of the margin-lever behaviour documented at DeepSeek and W&B.

Counterexamples

  • mistral-ai · Jul 2026 — Still no published cached-input pricing anywhere on La Plateforme — zero caching mentions in the corpus file — 23 months after Anthropic shipped prompt caching in August 2024. Pure per-token with no reuse discount, and Lago is its billing vendor, so this is not a metering-capability limit.
  • cohere · Jul 2026 — Still charges the full input-token rate regardless of prompt reuse; zero caching mentions in the corpus file. Together with Mistral, the two conspicuous holdouts among enterprise-positioned model vendors.
  • ai21-labs · Jul 2026 — Explicitly documents the absence: 'no batch discount, cached-input rate, or context-tier pricing published' — just two flat input/output pairs. The clearest statement in the corpus that a vendor has chosen to ship zero optimization levers, which removes the buyer's ability to engineer their own cost down.
  • cerebras · Jul 2026 — Sharpened counterexample: Cerebras now REPORTS caching without PRICING it. Its console carries a dedicated 'Cached-Usage' view showing cache hit rate and token breakdown, but no cached-input rate appears on the rate card — wafer-scale inference economics make the prefill saving structurally different from a multi-tenant GPU fleet's.
  • groq · Jul 2026 — The shallowest discount among active adopters: Kimi K2 lists at $1.00/$3.00 per 1M with cached input at $0.50 — only 2× (50% off), against the 10× that Anthropic, Moonshot and SambaNova now publish. Cache coverage also stays partial: not all models qualify and storage windows are constrained.

Trivia

  • Three vendors that share no infrastructure landed on exactly the same cache discount within nine days of each other: Moonshot's Kimi K3 at $0.30 against $3.00 input (2026-07-21), SambaNova's MiniMax-M2.7 at $0.06 against $0.60 (2026-07-23), and Anthropic's Opus 5 at $0.50 against $5.00 (2026-07-28). All three are precisely 10× — a 90%-off cached read now reads like a published industry constant rather than a per-vendor decision.

  • Prompt caching can cost you money. Anthropic charges 1.25× base input to write a 5-minute cache and 2× for a one-hour cache — $6.25/MTok to write against $5.00 to just send the tokens on Opus 5 (2026-07-28). Below roughly a 1-in-4 hit rate on the 5-minute cache the write premium exceeds the read saving, so the "discount" is a bet on your own traffic pattern, not a rebate.

  • xAI repriced grok-4.5 on 2026-07-21 without touching its headline: cached input fell 40% from $0.50 to $0.30 per 1M while uncached stayed at $2.00 in / $6.00 out. That is the cleanest example in the corpus of a vendor buying a specific workload — repeat-context agents — through the cache line, because the cache line is the one number a shopper comparing rate cards does not read first.

  • Zhipu AI (2026-07-22) is the first corpus vendor to name cache STORAGE as its own billable resource: a "Cached Input Storage" line appears on every model that supports caching, marked "Limited-time Free." A meter that exists, is documented, and currently bills zero is the standard corpus pattern for a price about to arrive.

  • 23 months after Anthropic shipped prompt caching in August 2024, Mistral and Cohere still publish no cached-input rate at all, and AI21 Labs documents the gap explicitly — "no batch discount, cached-input rate, or context-tier pricing published." Cerebras splits the difference in the strangest way: its console has a dedicated Cached-Usage dashboard reporting hit rates for a discount it does not offer.

See all pricing trivia

For buyers

If your application has a large, stable system prompt or RAG document context, cached-input pricing can cut input costs by 50-80%. Design your prompt architecture with caching in mind — keep the stable prefix at the front, variable parts after. But remember: cache investments are provider-specific; they raise switching costs.

For vendors

Cached-input pricing rewards your most loyal, highest-usage customers — those with established production pipelines with stable system prompts. The discount is a retention mechanism: once a customer has optimised their prompt architecture for your caching system, migration is expensive.

Outlook — what to watch

Expect caching to spread from frontier labs to more inference platforms. Baseten's February 2026 addition extended it to multi-tenant serving. Vendors without caching compete on price alone for the stable-prompt segment — that will push adoption. Watch for caching SLA tiers (guaranteed cache hit rates) as a premium feature.

Bottom line

Nine corpus vendors now publish cached-input pricing at 50-80% off standard input rates. Caching is a structural pricing tier that rewards stable-prompt workloads and raises switching costs.

FAQ

What is cached-input pricing in AI APIs?

A discounted per-token rate that applies when your input tokens match a previously stored prefix (cached context). Anthropic charges 75% off, OpenAI 50% off, Google 75% off, DeepSeek 74% off.

How much can I save with prompt caching?

If your application has a stable system prompt or document context, you can cut input costs by 50-80%. A 10k-token system prompt on Anthropic Claude, called 1,000 times, saves about $112 vs uncached.

Does caching work with RAG?

Yes — if your RAG pipeline prepends a fixed set of documents to every prompt, that prefix can be cached. Variable query context after the fixed prefix is still charged at the full input rate.

All trends