Ask
Launch

Fireworks ships Fire Pass and splits serverless into three named serving paths

Fireworks AI pricing

Fireworks launched Fire Pass, a promo-code pass with zero per-token charges for personal agentic coding, and now prices serverless on Standard, Priority and Fast paths.

Before

Serverless inference was framed as two quality-of-service tiers, Turbo and Priority, with per-model rates published in the docs.

After

Three named serving paths — Standard (default), Priority (service_tier parameter, higher reliability under peak traffic) and Fast (separate model ID, 100+ tokens/sec) — each with its own published input / cached-input / output rate. Fire Pass adds a zero-per-token, non-production pass on included open-weight models.

Proof of change

Fireworks AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured May 29, 2026 Captured Jul 22, 2026 · 54 days apart
docs
Captured May 29, 2026
Captured Jul 22, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 1 other page we compared
main
Captured May 29, 2026
Captured Jul 22, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • These captures are 54 days apart, so anything that changed and reverted in between would not appear here.
  • 4 pages had no counterpart in the earlier capture and are not shown.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Fireworks AI’s serverless pricing card now documents three serving paths rather than the earlier “Turbo + Priority” pairing. Standard is the default and needs no parameter. Priority is selected with a service_tier parameter, is prioritized above Standard traffic and is less likely to be load shed — it carries a 1.2x-1.5x uplift depending on the model (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority; OpenAI GPT OSS 120B is $0.15 versus $0.18). Fast is a separate model ID targeting 100+ tokens/sec of generated throughput at roughly double the Standard rate (Kimi K2.6 Fast at $2.00/1M input, GLM 5.2 Fast at $2.10).

Separately, Fireworks introduced Fire Pass — an experimental, promo-code-activated pass that removes per-token charges on a set of included open-weight models for use in agentic coding harnesses such as Claude Code, OpenCode, Cline, Kilo Code and OpenClaw. It issues a dedicated fpk_ API key that only works for those models, and is explicitly prohibited for production workloads. It is the first zero-marginal-cost packaging Fireworks has offered on an otherwise strictly metered platform.

The headline rate card is unchanged: H100 80 GB and H200 141 GB at $7.00/hr, B200 180 GB at $10.00/hr, B300 288 GB at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.

From Fireworks AI's pricing timeline
Three serving paths + Fire Pass; Series D and $1B ARR

Serverless inference is now documented as three named serving paths — Standard (default), Priority (higher reliability under peak traffic, set via service_tier) and Fast (100+ tokens/sec, selected by model ID) — replacing the earlier "Turbo + Priority" framing, with per-model input / cached-input / output rates published for each. Fireworks also shipped Fire Pass, an experimental promo-code pass that removes per-token charges on included open-weight models for personal agentic coding, and the site banner now announces a Series D and $1B ARR. Headline rate card (H100/H200 $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%) is unchanged.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity