Ask
Price change

Fireworks raises its US-only Serverless premium from 10% to 50%

Fireworks AI pricing

Fireworks now documents a 50% premium over base serverless prices for newly-launched US-only Serverless models, up from the 10% premium it published in July 2026; Kimi K3 US and the exempted GLM 5.2 Fast US keep their prior pricing.

Before

US-only Serverless routing carried a flat 10% premium over base model serverless prices, with one published exception (GLM 5.2 Fast US, no premium).

After

Beginning September 1, 2026, newly-launched US-only Serverless models carry a 50% premium over base model serverless prices; Kimi K3 US keeps its original (lower) premium and GLM 5.2 Fast US remains exempt.

Fireworks’ Serverless Pricing and US-only Serverless docs pages now read: “Beginning September 1, 2026, launched US-only Serverless models are priced at a 50% premium to the base model serverless prices. Kimi K3 US already includes this premium, while GLM 5.2 Fast US is an exception and matches global GLM 5.2 Fast pricing.” That replaces the flat-10%-on-any-model framing Fireworks had published since July 2026.

The publicly listed rate for Kimi K3 US ($3.30 / $0.33 / $16.50 against Kimi K3’s $3.00 / $0.30 / $15.00 Standard) is still exactly 10% above base, not 50% — so Fireworks’ own docs currently describe existing US-only rows and the new pricing rule inconsistently.

The available US-only model catalog also expanded alongside the policy change: US model IDs now exist for Kimi K3, DeepSeek V4 Flash (0731), GLM 5.2, GLM 5.2 Fast, GLM 5.3, and GLM 5.3 Flash — up from just Kimi K3 US and GLM 5.2 Fast US previously documented with individually priced rows.

From Fireworks AI's pricing timeline
On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%

The September 1, 2026 on-demand GPU price increase (announced 2026-08-12) is confirmed live: H100/H200 now $8.00/hr, B200 $13.00/hr, B300 $15.00/hr, GB300 $20.00/hr. Default self-serve GPU quotas also changed, scoped to the GLOBAL multi-region placement: H100/H200/B200/B300 doubled from 8 to 16 each, while A100's default quota dropped from 8 to 0 with no rate ever published for it. Separately, Fireworks' docs now state that, beginning September 1, 2026, newly-launched US-only Serverless models carry a 50% premium over base serverless pricing — up from the flat 10% premium published since July 2026 — while Kimi K3 US and the already-exempt GLM 5.2 Fast US keep their prior, lower pricing, so the docs describe an existing priced row and the new policy inconsistently.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity