Fireworks ships Fire Pass and splits serverless into three named serving paths
Fireworks launched Fire Pass, a promo-code pass with zero per-token charges for personal agentic coding, and now prices serverless on Standard, Priority and Fast paths.
Serverless inference was framed as two quality-of-service tiers, Turbo and Priority, with per-model rates published in the docs.
Three named serving paths — Standard (default), Priority (service_tier parameter, higher reliability under peak traffic) and Fast (separate model ID, 100+ tokens/sec) — each with its own published input / cached-input / output rate. Fire Pass adds a zero-per-token, non-production pass on included open-weight models.
Fireworks AI’s serverless pricing card now documents three serving paths rather than the earlier “Turbo + Priority” pairing. Standard is the default and needs no parameter. Priority is selected with a service_tier parameter, is prioritized above Standard traffic and is less likely to be load shed — it carries a 1.2x-1.5x uplift depending on the model (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority; OpenAI GPT OSS 120B is $0.15 versus $0.18). Fast is a separate model ID targeting 100+ tokens/sec of generated throughput at roughly double the Standard rate (Kimi K2.6 Fast at $2.00/1M input, GLM 5.2 Fast at $2.10).
Separately, Fireworks introduced Fire Pass — an experimental, promo-code-activated pass that removes per-token charges on a set of included open-weight models for use in agentic coding harnesses such as Claude Code, OpenCode, Cline, Kilo Code and OpenClaw. It issues a dedicated fpk_ API key that only works for those models, and is explicitly prohibited for production workloads. It is the first zero-marginal-cost packaging Fireworks has offered on an otherwise strictly metered platform.
The headline rate card is unchanged: H100 80 GB and H200 141 GB at $7.00/hr, B200 180 GB at $10.00/hr, B300 288 GB at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.
Serverless inference is now documented as three named serving paths — Standard (default), Priority (higher reliability under peak traffic, set via service_tier) and Fast (100+ tokens/sec, selected by model ID) — replacing the earlier "Turbo + Priority" framing, with per-model input / cached-input / output rates published for each. Fireworks also shipped Fire Pass, an experimental promo-code pass that removes per-token charges on included open-weight models for personal agentic coding, and the site banner now announces a Series D and $1B ARR. Headline rate card (H100/H200 $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%) is unchanged.