Fireworks ships Fire Pass and splits serverless into three named serving paths
Fireworks launched Fire Pass, a promo-code pass with zero per-token charges for personal agentic coding, and now prices serverless on Standard, Priority and Fast paths.
Serverless inference was framed as two quality-of-service tiers, Turbo and Priority, with per-model rates published in the docs.
Three named serving paths — Standard (default), Priority (service_tier parameter, higher reliability under peak traffic) and Fast (separate model ID, 100+ tokens/sec) — each with its own published input / cached-input / output rate. Fire Pass adds a zero-per-token, non-production pass on included open-weight models.
Proof of change
Fireworks AI's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 1 other page we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
- These captures are 54 days apart, so anything that changed and reverted in between would not appear here.
- 4 pages had no counterpart in the earlier capture and are not shown.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Fireworks AI’s serverless pricing card now documents three serving paths rather than the earlier “Turbo + Priority” pairing. Standard is the default and needs no parameter. Priority is selected with a service_tier parameter, is prioritized above Standard traffic and is less likely to be load shed — it carries a 1.2x-1.5x uplift depending on the model (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority; OpenAI GPT OSS 120B is $0.15 versus $0.18). Fast is a separate model ID targeting 100+ tokens/sec of generated throughput at roughly double the Standard rate (Kimi K2.6 Fast at $2.00/1M input, GLM 5.2 Fast at $2.10).
Separately, Fireworks introduced Fire Pass — an experimental, promo-code-activated pass that removes per-token charges on a set of included open-weight models for use in agentic coding harnesses such as Claude Code, OpenCode, Cline, Kilo Code and OpenClaw. It issues a dedicated fpk_ API key that only works for those models, and is explicitly prohibited for production workloads. It is the first zero-marginal-cost packaging Fireworks has offered on an otherwise strictly metered platform.
The headline rate card is unchanged: H100 80 GB and H200 141 GB at $7.00/hr, B200 180 GB at $10.00/hr, B300 288 GB at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.
Serverless inference is now documented as three named serving paths — Standard (default), Priority (higher reliability under peak traffic, set via service_tier) and Fast (100+ tokens/sec, selected by model ID) — replacing the earlier "Turbo + Priority" framing, with per-model input / cached-input / output rates published for each. Fireworks also shipped Fire Pass, an experimental promo-code pass that removes per-token charges on included open-weight models for personal agentic coding, and the site banner now announces a Series D and $1B ARR. Headline rate card (H100/H200 $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%) is unchanged.