Ask
Launch

Fireworks launches Reserved Throughput, a paid SLA capacity commit

Fireworks AI pricing

Fireworks added Reserved Throughput, a sales-led dollar-per-minute capacity reservation that guarantees availability and uptime on Serverless inference for Kimi K3, Kimi K3 Fast, DeepSeek V4 Flash (0731), and GLM 5.3.

Before

Serverless inference had no SLA-backed capacity option; all traffic ran on standard adaptive rate limits with no availability guarantee.

After

Reserved Throughput lets customers pre-purchase a dollar-per-minute reservation, SLA-backed up to the reservation level, with overage billed at standard Serverless rates (no SLA) and no rollover of unused capacity.

Fireworks shipped a new docs page and pricing mechanic: “Reserved throughput lets you pre-purchase throughput that guarantees availability and uptime up to your reservation level.” It targets customers hitting Serverless rate limits — the page’s own pitch is “If you are experiencing rate limits on Serverless, you should consider Reserved Throughput.”

The commercial mechanics are new to the platform: a customer buys a reserved dollar-per-minute amount (e.g. $10/minute); eligible usage is priced and deducted from that amount every minute; usage above the reservation in a given minute bills at standard Serverless rates with no SLA; and unused reservation is use-it-or-lose-it — it does not roll over. Fireworks’ own worked example sizes a Kimi K3 workload (5,000 uncached input, 50,000 cached input, 200 output tokens per request, p95 5 QPS) at $9.90/minute, rounding up to a $10/minute reservation.

Access is sales-led only (“Contact us to purchase reserved throughput or learn more”), and it is currently available on four models: Kimi K3, Kimi K3 Fast, DeepSeek V4 Flash (0731), and GLM 5.3 — the same four rows that newly carry a “Reserved Throughput” checkmark on the Serverless Pricing table. Reallocating a reservation between models is also not yet self-service; it requires contacting the account team.

From Fireworks AI's pricing timeline
On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%

The September 1, 2026 on-demand GPU price increase (announced 2026-08-12) is confirmed live: H100/H200 now $8.00/hr, B200 $13.00/hr, B300 $15.00/hr, GB300 $20.00/hr. Default self-serve GPU quotas also changed, scoped to the GLOBAL multi-region placement: H100/H200/B200/B300 doubled from 8 to 16 each, while A100's default quota dropped from 8 to 0 with no rate ever published for it. Separately, Fireworks' docs now state that, beginning September 1, 2026, newly-launched US-only Serverless models carry a 50% premium over base serverless pricing — up from the flat 10% premium published since July 2026 — while Kimi K3 US and the already-exempt GLM 5.2 Fast US keep their prior, lower pricing, so the docs describe an existing priced row and the new policy inconsistently.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity