Ask
Sharpens 17 companies · First observed August 2024 · Updated August 2026 Explore in the graph

The per-token list price is becoming a base rate times named modifiers

Quick answer

Inference vendors have converged on two standard discounts — roughly 50% off for latency-tolerant batch jobs and a reduced rate for cached input. If part of your workload is async or shares a stable prompt prefix, these are the highest-leverage cuts on a token bill.

~50% off for batch / async workloads

What's happening — and why

What's happening: token APIs now routinely offer two big discounts — roughly half price for 'batch' jobs you're willing to wait on, and a reduced rate for cached (repeated) input such as a fixed system prompt or RAG context.

Why: latency-tolerant and repetitive work is cheaper for the vendor to serve — it can be scheduled onto idle capacity or skip recomputation. Pricing it lower sorts that load onto cheaper infrastructure and rewards buyers for flexibility, all without touching the headline real-time rate.

How it works

real-time cached input batch (async) 100% −up to 80% −50% (batch)
Price falls as you trade latency: full rate → cached input → ~50%-off batch.

Evidence over time

25 supporting · 6 counter — hover or tap a point for detail, click to jump to the row.

supports ↑ challenges ↓ 2024 2025 2026
supporting evidence counterexample

Evidence

Company Date What happened
OpenAI May 2026 GPT-5.x line ships with Batch API (50% off) and prompt caching on the reduced cached-input rate — the standard pair on the flagship API.
Google May 2025 Gemini 2.5 added Priority (1.8×) and Flex/Batch (0.5×) as named SKUs — explicit latency-tier price discrimination plus the 50% batch rate.
Fireworks AI Mar 2025 Batch API at a flat 50% discount across all models.
Groq May 2025 Batch API plus cached-input discounts launched together.
Anthropic Aug 2024 Prompt caching cut input cost by up to 80%; batch API also offered at 50%.
Baseten Feb 2026 Cached-input pricing added to multi-tenant Model APIs.
Fireworks AI Nov 2024 Cached-input discount shipped alongside Turbo / Priority latency tiers.
Mistral AI May 2026 Batch processing earns a 50% discount on per-token rates.
OpenRouter Jun 2025 Multi-model routing marketplace that passes through upstream batch/cache discounts — extending the standard playbook to routing infrastructure without re-inventing the mechanics.
Anthropic Jun 2026 Fast mode (research preview) prices the same model higher for lower latency — Opus 4.8 at $10/$50 vs standard $5/$25 (2× premium), Opus 4.6/4.7 at $30/$150 — the surcharge-for-speed half of the latency axis, alongside its existing 50% Batch and prompt-caching discounts. Not available on AWS or Batch.
DeepInfra Jun 2026 Added a per-request Service Tier: Standard (1× base, default) or Priority (1.5× base, scheduled ahead of standard traffic for faster TTFT), set via service_tier: "priority". First self-serve per-request speed surcharge on a pure-usage per-token platform — the surcharge half of latency tiering reaching the multi-tenant inference layer.
DeepInfra Jul 2026 Completed the ladder three weeks after starting it: a third per-request Service Tier, Flex at 0.8× base — "lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability" — joining Standard (1×) and Priority (1.5×). Flex also became a filter and a per-model badge in the model directory. Headline per-token rates were untouched throughout, so the same model now spans 1.9× (0.8× to 1.5×) on one platform purely by scheduling choice. This is the corpus's only patience discount published as a multiplier on the primary rate card rather than as a separate batch endpoint. Verified in src/content/blueprint/deepinfra.mdx (pricing-page capture 2026-07-29).
Modal Jul 2026 Two implicit multipliers became published line items: region selection at 1.5–1.75× base prices and non-preemptible (guaranteed, non-interruptible) execution at 3× base prices, both applying across Starter, Team and Enterprise. The page now states outright that "headline per-second rates assume preemptible, in-default-region scheduling." 3× is the steepest published reliability/latency modifier in the corpus. Shipped alongside B300 at $0.001972/sec and a separately-metered Sandbox+Notebooks tier at roughly 3× standard function rates; all base rates and plan fees ($0 / $250 / Custom) unchanged. Verified in src/content/blueprint/modal.mdx.
Fireworks AI Jul 2026 Formalised three named serving paths on the same models, replacing the older "Turbo + Priority" framing: Standard (default, no parameter), Priority (higher reliability under peak traffic, roughly 1.25–1.5× Standard, set via service_tier) and Fast (100+ tokens/sec, roughly 2× Standard, selected by model ID — Kimi K2.6 Fast at $2.00/1M input vs Standard). The path is chosen two different ways depending on which axis you want, which is itself a sign the modifier layer is accreting rather than being designed.
Fireworks AI Jul 2026 Added the corpus's first GEOGRAPHY modifier on the same model: a flat 10% premium over base serverless pricing for any model routed US-only, documented alongside the Kimi K3 launch (plus a Fast variant and a US-only variant of it). That makes four independent axes on one rate card — model, serving path, US-only routing, and cached-vs-uncached input.
Together AI Jul 2026 Guaranteed throughput sold as its own SKU rather than a multiplier: Provisioned Throughput reserves dedicated capacity in throughput units billed per PTU-minute at $0.05/PTU-min on MiniMax M3 and GLM-5.2, "for buyers who prefer a fixed" cost. Shipped in the same release as a Dedicated Inference restructure and rate cuts — the discount and the guarantee arriving together.
Baseten Feb 2026 The reliability half of the same axis, sold as an enterprise add-on rather than a per-request flag: a Mission Critical SLA tier formalised at 99.95% uptime with 24/7 incident response, quote-based and typically tied to annual usage commitments. Baseten's own tradeoff table names the gap — SOC 2 Type II and HIPAA ship on Basic, but there is no published SLA below Enterprise. Shipped alongside B200 at $0.16633/min and cached-input pricing on the multi-tenant Model APIs.
Voyage AI Jun 2026 Two modifiers on an embeddings rate card: a 33% Batch API discount with a 12-hour completion window (note: 33%, not the industry-standard 50% — the batch number is less universal outside LLM inference), and throughput that "scales with cumulative spend rather than a paid plan" — $100 billed unlocks 2× rate limits, $1,000 unlocks 3×. Rate limits priced by loyalty rather than by plan is unique in the corpus.
You.com Jul 2026 The widest single-endpoint price spread in the corpus, sold purely on declared effort: research_effort tiers at lite $12, standard $50, deep $100, exhaustive $450 and Frontier $1,200 per 1,000 calls — a 100× range on one API. The July 2026 change replaced Frontier's open-ended ">$2,000.00 /1k calls" label with the concrete $1,200 figure. Effort, like latency, is a buyer-declared quality knob priced as a multiplier on the same endpoint.
Replit AI May 2026 Effort as the meter itself rather than a multiplier on one: Agent and Assistant requests each become a checkpoint "priced on the time and compute it actually consumed, not a flat per-task fee" — simple changes typically under $0.25, a full feature build $1–$3, drawn against included monthly credits then pay-as-you-go. Replaced an earlier flat $0.25-per-checkpoint model, i.e. moved FROM a flat unit TO a consumption-weighted one.
Gladia Jul 2026 Latency-as-quality reaching a speech API's mid tier, and disclosed for the first time: a new plan-comparison table makes explicit that Growth (not just Enterprise) carries a 99.9% uptime SLA and a priority processing queue, while Starter has neither. "This is the first time Gladia's pricing page has stated a numeric uptime SLA for Growth." No per-hour rate changed — Starter async $0.61/hr and real-time $0.75/hr, Growth as low as $0.20/$0.25 with an upfront commitment. The discount and the priority queue are bundled into the same commitment gate.
Luma AI Jul 2026 The patience discount in a consumer credit product: the top Ultra tier ($300/mo) adds an unmetered relaxed (slower) mode — capacity traded for latency, exactly Midjourney's relax mechanic inside a credit plan. Worth noting what is NOT a latency signal here: the plan cards' 4×/15× "usage multipliers" are simply the credit totals restated against Plus (40,000 is 4× of 10,000; 150,000 is 15×), not a price modifier.
DeepSeek Aug 2026 A NEW AXIS: time of day, and the first modifier in this trend the buyer cannot select. DeepSeek's docs disclosed the policy on 2026-08-06 as a footnote (2x the regular rate during 9:00-12:00 and 14:00-18:00 Beijing Time, no effective date) and published the card on 2026-08-14: effective 16:00 UTC on 2026-08-16, peak hours are 01:00-04:00 and 06:00-10:00 UTC and all other hours are off-peak at exactly half the peak rate. DeepSeek-V4-Flash cache-miss input goes from a flat $0.14 per 1M to $0.22 off-peak / $0.44 peak and output from $0.28 to $0.66 / $1.32; V4-Pro cache-miss input from $0.435 to $0.66 / $1.32, with the steepest relative jump on cache-hit input. The multiplier is exactly 2x peak-over-off-peak — inside the 'fast ~2x' band the corpus converged on — but it is set by the clock, not by a service_tier flag, a model ID or a region. Thirteenth bidirectional vendor, and the only one whose modifier a buyer avoids by rescheduling rather than by choosing.
Fireworks AI Aug 2026 The geography modifier escalated from 1.1x to 1.5x and moved from tokens to silicon, at the vendor that invented it thirteen days earlier. Region-restricted deployments — dedicated GPUs pinned to US-only or Europe-only infrastructure — are priced at a flat 1.5x the standard on-demand rate and gated behind a Contact Sales request rather than self-serve checkout, alongside GB300 288 GB entering the card at $18.00/hr as the highest published on-demand rate. This is the first geographic-routing premium Fireworks has published for dedicated GPU deployments and it parallels the 10% US-only Serverless premium logged on 2026-07-29 — so one rate card now prices data residency two ways at two magnitudes, and the premium is larger where the resource is more physical.
Modal Aug 2026 The pricing SURFACE itself becomes the gated tier — a fourth procurement shape alongside per-request flag, plan gate and enterprise SLA. Modal's July 29 announcement of an OpenAI-compatible, token-billed Shared API originally said it was covered by Starter's $30/month free-compute offer, implying access on any plan; the same post now states that Shared API token-based pricing is available to Team ($250/mo + compute) and Enterprise customers only. Starter and lower tiers keep access to the same hosted models — Kimi K3 among them — through a per-second Auto Endpoint. They lose the meter, not the model: the cheaper plan is billed on a different axis entirely. Modal has still published no per-token input/output rates for the Shared API on its pricing page or billing docs.

Counterexamples

  • Suno · May 2026 — Consumer credit tiers — no batch or cache discount.
  • Midjourney · Feb 2025 — Uses fast vs relax compute modes instead of cache/batch — latency tiering by queue, not caching.
  • Lambda Labs · Jun 2026 — GPU cloud with no batch API or caching discount — the playbook is specific to token/inference APIs, not raw compute rental.
  • SambaNova · Jul 2026 — The multiplier half's clearest boundary case. SambaNova's whole pitch is speed — custom RDU silicon sold on tokens-per-second — and it DID add cached-input pricing in July 2026 (MiniMax-M2.7 at $0.06/1M cached, "the first model on the card to carry cached-input pricing"), so it adopted the discount half. But its six-model rate card carries no priority, fast, flex or service-tier multiplier of any kind: one rate per model, take it or leave it. A vendor whose differentiator IS latency declining to sell latency as a tier.
  • Zhipu AI · Jul 2026 — Same boundary: Zhipu publishes aggressive cached-input rates ($0.01–$0.45 per 1M, with cached-input storage listed "Limited-time Free") and three free Flash models — heavy use of the DISCOUNT side — while its z.ai rate card has one price per model with no priority, fast or guaranteed tier. It reprices the model instead of tiering the service: GLM Coding Plan moved to $18/$72/$160 a month and a premium GLM-5 API tier was added on 2026-07-22, i.e. a new SKU rather than a multiplier.
  • Sarvam AI · Jul 2026 — Third boundary case, and the one with the most obvious commercial reason to add a multiplier: Sarvam publishes per-model cached rates (Sarvam-30B ₹2.5 in / ₹1.5 cached / ₹10 out; Sarvam-105B ₹4 / ₹2.5 / ₹16) and sells prepaid plans that differ ONLY by request-per-minute limit — Starter free / Pro ₹10,000 / Business ₹50,000 at 60 / 200 / 1,000 req-min. That is throughput sold by plan tier, not by per-request multiplier: the same demand DeepInfra prices at 1.5× per call, Sarvam prices as a prepaid commitment. Its 2026-07-23 change raised prepaid bonus credits and moved "Most Popular" to Business rather than adding a service tier.

Trivia

  • Anthropic's August 2024 prompt-caching launch — cutting repeated-input costs by up to 80% — was the largest single-day effective price reduction in the corpus that was not accompanied by a new model release. It created a new cost category (cached vs uncached input) without changing the headline per-token rate, meaning vendors could simultaneously advertise stable pricing while offering dramatically lower effective costs to buyers with stable system prompts.

  • The ~50% batch discount has converged across at least 13 corpus vendors (Anthropic, OpenAI, Google, Mistral, Fireworks, Groq, Together, Replicate, and others) to the same approximate number — a rare case of industry-wide price coordination on a structural discount. The consistency suggests the discount reflects a genuine cost difference (asynchronous batching reduces per-token GPU utilisation variance) rather than arbitrary marketing.

  • Google's addition of a 1.8× Priority premium on Gemini 2.5 (May 2025) is the corpus's first example of latency tiering working in both directions simultaneously: a discount for tolerance (Flex/Batch at 0.5×) and a premium for impatience (Priority at 1.8×) on the same model. The bidirectional approach makes latency an explicit price axis rather than an implicit quality attribute — a structural move that no other corpus vendor has fully replicated.

  • Anthropic's June 2026 Fast mode inverts the usual generational price order: the newer Opus 4.8 costs $10/$50 in Fast mode while the older Opus 4.6/4.7 cost $30/$150 — 3× more — to run fast. Premium-speed pricing tracks inference efficiency, not recency, so the prior-generation flagship becomes the *expensive* one the moment latency is the SKU. It is the corpus's second bidirectional latency example after Google, and the first where the speed surcharge is steeper on the model you'd expect to be cheaper.

  • DeepInfra's June 2026 Priority tier (1.5× base) lands between Google's 1.8× and Anthropic's 2× speed premiums — the three bidirectional-latency vendors have independently converged on a 1.5×–2× surcharge band for priority scheduling, a tighter cluster than the discount side, where batch settled on a single ~50% number. The speed surcharge is the only price lever in the corpus that a buyer sets per individual request (service_tier: "priority") rather than per plan or per model.

  • DeepInfra built a complete three-step cost/latency ladder in three weeks without touching a single headline rate: Priority at 1.5× base on 2026-06-30, then Flex at 0.8× base on 2026-07-21, around an unchanged Standard at 1×. That 0.8× Flex is the corpus's only published patience DISCOUNT expressed as a multiplier on the same rate card rather than as a separate batch endpoint — and the Flex-to-Priority spread means the same model on the same platform ranges 1.9× in price with no model change at all.

  • Modal's 2026-07-14 update is the rare case of a vendor pricing MORE transparently by admitting it was already charging more: region selection at 1.5–1.75× and non-preemptible execution at 3× base had both been implicit until they appeared as published line items. Its own page now warns that "headline per-second rates assume preemptible, in-default-region scheduling." The 3× non-preemptible multiplier is the steepest published latency/reliability modifier in the corpus.

  • You.com prices the widest single-endpoint spread in the corpus purely by declared effort: research_effort runs lite $12, standard $50, deep $100, exhaustive $450 and Frontier $1,200 per 1,000 calls — a 100× range on one API, with the July 2026 change replacing Frontier's open-ended ">$2,000/1k calls" label with a concrete $1,200. Fireworks meanwhile added the corpus's first GEOGRAPHY modifier on 2026-07-29: a flat 10% premium for US-only serverless routing, a fourth axis on top of Standard/Priority/Fast.

See all pricing trivia

For buyers

If a meaningful share of your workload is asynchronous, batch roughly halves spend with no model change; if your prompts share a long stable prefix (system prompts, RAG context), caching compounds the saving. These are the highest-leverage, lowest-effort moves on a token bill.

For vendors

The playbook needs a batch queue with a relaxed SLA and a prompt-cache keyed on prefix hashes, each priced as its own line. The discounts are a segmentation tool — they sort latency-tolerant load onto cheaper infra without dropping your headline rate.

Outlook — what to watch

Expect these to become table stakes and to deepen: longer cache TTLs, automatic prompt-prefix caching, and tiered batch SLAs (1-hour vs 24-hour). The next frontier is priority/express pricing in the other direction — paying a premium for guaranteed low latency — turning latency into a full price axis.

Bottom line

Inference vendors have converged on ~50%-off batch and cached-input discounts as a de-facto standard. Anthropic, Mistral, Fireworks and Groq all land near the same numbers.

FAQ

How can I cut my LLM API bill without changing models?

Use batch processing for anything asynchronous (≈50% off) and prompt caching for repeated context like system prompts or RAG (a further large cut on input). Both are vendor-native and need no model change.

What is prompt caching?

A discount on input tokens that repeat across requests — the vendor caches a stable prefix (e.g. your system prompt) and charges a fraction of the normal rate to reuse it. Anthropic's cut input cost by up to 80%.

Which vendors offer batch and cache discounts?

It's now near-standard for token APIs — Anthropic, Mistral, Fireworks, Groq and Baseten all offer batch (~50%) and/or cached-input pricing. Consumer credit apps generally don't.

All trends