Fireworks launches Reserved Throughput, a paid SLA capacity commit
Fireworks added Reserved Throughput, a sales-led dollar-per-minute capacity reservation that guarantees availability and uptime on Serverless inference for Kimi K3, Kimi K3 Fast, DeepSeek V4 Flash (0731), and GLM 5.3.
Serverless inference had no SLA-backed capacity option; all traffic ran on standard adaptive rate limits with no availability guarantee.
Reserved Throughput lets customers pre-purchase a dollar-per-minute reservation, SLA-backed up to the reservation level, with overage billed at standard Serverless rates (no SLA) and no rollover of unused capacity.
Fireworks shipped a new docs page and pricing mechanic: “Reserved throughput lets you pre-purchase throughput that guarantees availability and uptime up to your reservation level.” It targets customers hitting Serverless rate limits — the page’s own pitch is “If you are experiencing rate limits on Serverless, you should consider Reserved Throughput.”
The commercial mechanics are new to the platform: a customer buys a reserved dollar-per-minute amount (e.g. $10/minute); eligible usage is priced and deducted from that amount every minute; usage above the reservation in a given minute bills at standard Serverless rates with no SLA; and unused reservation is use-it-or-lose-it — it does not roll over. Fireworks’ own worked example sizes a Kimi K3 workload (5,000 uncached input, 50,000 cached input, 200 output tokens per request, p95 5 QPS) at $9.90/minute, rounding up to a $10/minute reservation.
Access is sales-led only (“Contact us to purchase reserved throughput or learn more”), and it is currently available on four models: Kimi K3, Kimi K3 Fast, DeepSeek V4 Flash (0731), and GLM 5.3 — the same four rows that newly carry a “Reserved Throughput” checkmark on the Serverless Pricing table. Reallocating a reservation between models is also not yet self-service; it requires contacting the account team.
The September 1, 2026 on-demand GPU price increase (announced 2026-08-12) is confirmed live: H100/H200 now $8.00/hr, B200 $13.00/hr, B300 $15.00/hr, GB300 $20.00/hr. Default self-serve GPU quotas also changed, scoped to the GLOBAL multi-region placement: H100/H200/B200/B300 doubled from 8 to 16 each, while A100's default quota dropped from 8 to 0 with no rate ever published for it. Separately, Fireworks' docs now state that, beginning September 1, 2026, newly-launched US-only Serverless models carry a 50% premium over base serverless pricing — up from the flat 10% premium published since July 2026 — while Kimi K3 US and the already-exempt GLM 5.2 Fast US keep their prior, lower pricing, so the docs describe an existing priced row and the new policy inconsistently.