Launch

Fireworks ships Fire Pass and splits serverless into three named serving paths

Fireworks AI pricing

Fireworks launched Fire Pass, a promo-code pass with zero per-token charges for personal agentic coding, and now prices serverless on Standard, Priority and Fast paths.

Before

Serverless inference was framed as two quality-of-service tiers, Turbo and Priority, with per-model rates published in the docs.

After

Three named serving paths — Standard (default), Priority (service_tier parameter, higher reliability under peak traffic) and Fast (separate model ID, 100+ tokens/sec) — each with its own published input / cached-input / output rate. Fire Pass adds a zero-per-token, non-production pass on included open-weight models.

Fireworks AI’s serverless pricing card now documents three serving paths rather than the earlier “Turbo + Priority” pairing. Standard is the default and needs no parameter. Priority is selected with a service_tier parameter, is prioritized above Standard traffic and is less likely to be load shed — it carries a 1.2x-1.5x uplift depending on the model (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority; OpenAI GPT OSS 120B is $0.15 versus $0.18). Fast is a separate model ID targeting 100+ tokens/sec of generated throughput at roughly double the Standard rate (Kimi K2.6 Fast at $2.00/1M input, GLM 5.2 Fast at $2.10).

Separately, Fireworks introduced Fire Pass — an experimental, promo-code-activated pass that removes per-token charges on a set of included open-weight models for use in agentic coding harnesses such as Claude Code, OpenCode, Cline, Kilo Code and OpenClaw. It issues a dedicated fpk_ API key that only works for those models, and is explicitly prohibited for production workloads. It is the first zero-marginal-cost packaging Fireworks has offered on an otherwise strictly metered platform.

The headline rate card is unchanged: H100 80 GB and H200 141 GB at $7.00/hr, B200 180 GB at $10.00/hr, B300 288 GB at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.

From Fireworks AI's pricing timeline
Three serving paths + Fire Pass; Series D and $1B ARR

Serverless inference is now documented as three named serving paths — Standard (default), Priority (higher reliability under peak traffic, set via service_tier) and Fast (100+ tokens/sec, selected by model ID) — replacing the earlier "Turbo + Priority" framing, with per-model input / cached-input / output rates published for each. Fireworks also shipped Fire Pass, an experimental promo-code pass that removes per-token charges on included open-weight models for personal agentic coding, and the site banner now announces a Series D and $1B ARR. Headline rate card (H100/H200 $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%) is unchanged.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints for popular open-weight models (Llama, DeepSeek, Mixtral, Qwen) plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Jul 2026
    Three serving paths + Fire Pass; Series D and $1B ARR
  2. Feb 2026
    Embeddings Pricing by Parameter Size
  3. Aug 2025
    B200 and B300 GPU Availability — Frontier Pricing
  4. Mar 2025
    Batch API at 50% Discount Across All Models
  5. Nov 2024
    Turbo + Priority Tiers + Cached Input Discount
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity