Ask
All companies
technology

Fireworks AI pricing

fireworks.ai facts checked analysis reviewed
Estimate your Fireworks AI cost — model your usage, see overages, and find the cheapest plan. Open calculator →
Quick summary
Product
Generative AI inference platform — serverless per-token, on-demand GPU, fine-tuning, batch API, serverless training API
Industry
technology
Commits
Available (annual)
In this page
AI Summary
  • Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.
  • Serverless inference runs on three serving paths — Standard (default), Priority (higher reliability during peak traffic, roughly 1.25–1.5× the Standard rate) and Fast (100+ tokens/sec, roughly 2× Standard) — each with its own published per-model input / cached-input / output rate; batch inference bills at 50% of serverless on both input and output. A newly documented US-only Serverless option adds a fourth axis: a flat 10% premium over base serverless pricing for models routed to US-only infrastructure, with one published exception (GLM 5.2 Fast US carries no premium and matches the global rate).
  • Fine-tuning is priced per 1M training tokens by model size: <16B at $0.50 LoRA SFT / $1.00 LoRA DPO / $1.00 full SFT / $2.00 full DPO; 16–80B at $3.00–$12.00; 80–300B at $6.00–$24.00; >300B at $10.00–$40.00. A separate Serverless Training API (added 2026-07-29) instead meters LoRA training continuously per token — prefill, cached prefill, sample, and train — on a shared, always-on trainer pool with no provisioning or idle cost, for Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3.
  • Embeddings priced by parameter size: <150M params at $0.008/1M tokens, 150–350M at $0.016/1M, and Qwen3 8B at $0.10/1M — undercutting OpenAI's text-embedding-3 rates by 80–90% for the smaller models.
  • Free access comes in two forms: a $1 trial credit (one of the smallest in the market) and Fire Pass, an experimental promo-code pass launched in July 2026 that removes per-token charges entirely on included open-weight models for personal agentic coding, issues a separate fpk_ API key, and is prohibited for production workloads. Paid self-serve accounts run on prepaid credits with spending tiers that cap the monthly budget at $50 (Tier 1), $500 (Tier 2), $5,000 (Tier 3) and $50,000 (Tier 4); Enterprise accounts are exempt from those caps and are quote-based.
  • Founded 2022 by Lin Qiao (ex-Meta PyTorch lead), Dmytro Ivchenko, and Pawel Garbacki; raised Series B in 2024 led by Sequoia at $552M valuation, with Series C reported in 2025 at $4B+ valuation, and announced a Series D and $1B ARR milestone on its own site in July 2026.
Pricing summary
Fireworks AI 2026 — Per-token serverless + per-hour dedicated
$1 free credits; Kimi K3 flagship model; Standard / Priority / Fast serving paths; batch at 50%; new Serverless Training API
Free trial
$1 credit
Evaluating Fireworks for proof-of-value
Contact sales
Enterprise
Custom
Sustained high-volume production workloads
Price rising Sep 1
Dedicated GPU
From $7.00 /hr (H100, through Aug 31)
Single-tenant per-GPU-second deployments
Fine-tuning
From $0.50 /1M tokens
LoRA + full-parameter SFT and DPO
Experimental
Fire Pass
Promo code
Personal agentic coding, non-production only
New 2026-07-29
Serverless Training API
From $0.66 /1M prefill tokens
Tinker-compatible LoRA training, shared trainer pool
No monthly fee; pay-as-you-go on every SKU, billed against prepaid credits. Batch inference bills at 50% of serverless on both input and output, and cached input tokens are priced at 50% of the input rate by default (headline models are specified lower still). Reinforcement fine-tuning is billed per GPU hour at the same rate as on-demand deployments. On-demand GPU rates are increasing across the board on Sep 1, 2026 — see Pricing by product for the full before/after.

About

Fireworks AI is a Redwood City-based generative AI infrastructure company founded in October 2022 by Lin Qiao (former PyTorch team lead at Meta), Dmytro Ivchenko, and Pawel Garbacki. The product is a high-performance inference platform optimized for serving open-source and customer-tuned models behind production endpoints — combining a serverless per-token API for popular open-weight models with on-demand dedicated GPU deployments and a fine-tuning service. The runtime is built on Fireworks’ proprietary FireAttention kernels, FireOptimizer auto-tuning, and speculative-decoding pipelines designed to extract higher throughput from each GPU than open-source serving frameworks deliver.

By 2026 Fireworks serves Cursor, Notion, Doordash, Quora, Upwork, and roughly a thousand other paying customers across enterprise AI infrastructure (RAG systems, customer-facing assistants, agentic workflows) and developer-tooling startups serving sub-second inference SLAs. The company raised a $52M Series B in July 2024 led by Sequoia Capital at a $552M post-money valuation with Bessemer, Benchmark, and NVIDIA participation; a Series C reported in 2025 brought valuation past $4B. As of July 2026 the site’s own top banner announces a Series D and a $1B ARR milestone, and on 2026-07-29 the same banner promoted Kimi K3, a new flagship model added to the serverless rate card alongside a new Serverless Training API product for pay-per-token LoRA training.

Fireworks competes directly with Together AI, Baseten, Replicate, and Groq for the managed-inference market, and with first-party providers (OpenAI, Anthropic, Cohere) for general-purpose API customers. Its differentiation is the combination of PyTorch-team founder credibility, aggressive per-hour H100 pricing ($7.00/hr — among the lowest in the market), and a granular fine-tuning rate card that lets cost-conscious teams choose between LoRA (cheap, fast) and full-parameter (expensive, higher quality) workflows.


Pricing summary : How Fireworks AI’s serverless + dedicated + fine-tuning stack works

Fireworks bills on five parallel meters and no subscription. Serverless inference charges per 1M tokens across three separately-priced dimensions — input, cached input, and output — on three serving paths: Standard (the default, no parameter needed), Priority (service_tier: "priority", prioritized above Standard traffic and less likely to be load shed, at a higher price point), and Fast (a separate model ID targeting 100+ tokens/sec on the same model, also at a higher price point). As of 2026-07-29 the headline rate card is led by the newly launched Kimi K3 ($3.00 / $0.30 / $15.00 Standard, $4.50 / $0.45 / $22.50 Fast), with a Kimi K3 US variant carrying a flat 10% US-only Serverless premium. Batch inference bills at 50% of serverless on both input and output, and cached input tokens are priced at 50% of input by default for text and vision language models “unless otherwise specified” — the headline models on the docs card are specified far below that. On-demand deployments are single-tenant GPUs billed per GPU-second with no start-up charge, published at $7.00/hr for both H100 80 GB and H200 141 GB, $10.00/hr for B200 180 GB, $12.00/hr for B300 288 GB, and $18.00/hr for GB300 288 GB — through August 31, 2026. New as of this capture: Fireworks has published an across-the-board on-demand price increase effective September 1, 2026 — H100/H200 rising to $8.00/hr (+14%), B200 to $13.00/hr (+30%), B300 to $15.00/hr (+25%), and GB300 to $20.00/hr (+11%). This is the first repricing of an existing on-demand SKU in the tracked history; every prior on-demand change (B200/B300 in 2025, GB300 in 2026-08) was a new-tier addition rather than a rate change on an already-published GPU. A second new mechanic on the on-demand card prices geography directly: region-restricted deployments (US, Europe) are priced at a flat 1.5x the standard on-demand rate and require a Contact Sales request. Fine-tuning is priced per 1M training tokens by base-model size and method (LoRA vs full-parameter, SFT vs DPO), and reinforcement fine-tuning is instead billed per GPU hour at the same rate as on-demand deployments. New alongside it, the Serverless Training API (Tinker-compatible) prices LoRA training on a shared, always-on trainer pool per token prefilled/sampled/trained — $0.66–$10.87 per 1M prefill tokens across Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3 — with no provisioning step and no idle cost. Embeddings bill on input tokens only, at $0.008–$0.10 per 1M by base-model parameter count.

Onboarding is $1 in free credits, and self-serve accounts run on prepaid credits inside four spending tiers gated by cumulative spend or added credits: Tier 1 (valid payment method), Tier 2 ($50), Tier 3 ($500), Tier 4 ($5,000), and Unlimited on request. As of this capture, Fireworks’ docs clarify that the long-published $50/$500/$5,000/$50,000 monthly-spend ceilings apply specifically to legacy self-serve postpaid accounts; today’s default prepaid accounts instead set their own monthly spend limit independently of tier via firectl quota update monthly-spend-usd, while the tier itself still gates serverless TPM upper bounds and the training-GPU allocation. Enterprise accounts sit outside the tier system entirely — no spend cap, no spend-triggered pause, and the option to move from prepaid to post-paid billing. Alongside the metered SKUs, Fire Pass is an experimental, promo-code-activated pass that removes per-token charges on a set of included open-weight models for personal agentic coding, explicitly prohibited for production use — as of 2026-07-29 the granted model is Kimi K3 Fast (1M-token context), up from GLM 5.2 Fast (256k context). This mostly pure-usage architecture — token / GPU-second / training-token, all drawn against a prepaid credit balance — is one of the most granular rate cards in the AI infrastructure pricing landscape.

What makes this different: Fireworks prices reliability and speed as two separate, independently-purchasable axes on the same model. Priority buys queue precedence, Fast buys throughput, and each carries its own published per-model rate — so a buyer can pay for load-shed protection without paying for speed, or vice versa, rather than accepting a single bundled “premium tier”.


Pricing by product

Serverless inference (headline models, per 1M tokens)

Each cell reads input / cached input / output. A dash in the Priority column means Priority is not offered for that model; the docs price card is the source of truth for Priority availability. “Fast” rows are the same model on the high-speed serving path, selected by a different model ID. As of 2026-07-29 the docs also publish a US-only Serverless premium: endpoints pinned to US-only routing are priced at a flat 10% markup over the base model’s serverless prices (Kimi K3 US is the first model to carry its own priced row for this — $3.30 / $0.33 / $16.50 Standard against Kimi K3’s $3.00 / $0.30 / $15.00). As of this capture the docs note one published exception to that 10% rule: GLM 5.2 Fast US is priced identically to global GLM 5.2 Fast ($2.10 / $0.21 / $6.60) rather than carrying the premium.

ModelStandard (in / cached / out)Priority (in / cached / out)Key mechanics
Kimi K3$3.00 / $0.30 / $15.00$3.75 / $0.375 / $18.75New flagship model (2026-07-29); Priority is a flat 1.25× uplift
Kimi K3 Fast$4.50 / $0.45 / $22.501.5× the Standard row; no Priority path; this is the model behind the new Fire Pass grant
Kimi K3 US$3.30 / $0.33 / $16.50$4.125 / $0.4125 / $20.625Exactly 10% above Kimi K3 Standard/Priority — the published US-only Serverless premium
Kimi K2.7 Code$0.95 / $0.19 / $4.00$1.425 / $0.285 / $6.00Priority is a flat 1.5× uplift on every dimension
Kimi K2.7 Code Fast$1.90 / $0.38 / $8.00Exactly 2× the Standard row; no Priority path
Kimi K2.6$0.95 / $0.16 / $4.00$1.50 / $0.22 / $6.00Cheapest cached rate in the Kimi family
Kimi K2.6 Fast$2.00 / $0.30 / $8.00Speed premium priced into the model ID
DeepSeek V4 Pro$1.74 / $0.145 / $3.48$2.61 / $0.218 / $5.22Cached input at ~8% of the input rate
DeepSeek V4 Flash (0731)$0.22 / $0.007 / $0.66$0.275 / $0.00875 / $0.825Repriced as of this capture — the separate undated “DeepSeek V4 Flash” row (previously identical to this row) has been removed; the dated (0731) row is now the sole DeepSeek V4 Flash entry and its published rate moved input +57%, cached input −75%, output +136% versus the prior rate (see Pricing evolution)
GLM 5.2$1.40 / $0.14 / $4.40$1.75 / $0.18 / $5.50Priority uplift only 1.25× here
GLM 5.2 Fast$2.10 / $0.21 / $6.60The model behind Fire Pass (glm-5p2-fast)
GLM 5.2 Fast US$2.10 / $0.21 / $6.60New row; the one published exception to the 10% US-only Serverless premium
GLM 5.1$1.40 / $0.26 / $4.40$2.10 / $0.39 / $6.60Same Standard price as 5.2, worse cached rate
GLM 5.1 Fast$2.80 / $0.52 / $8.802× Standard
Qwen 3.7 Plus$0.40 / $0.08 / $1.60Standard-only
Qwen 3.8 Max$2.00 / $0.25 / $6.00New row as of this capture; Standard-only
MiniMax M3$0.30 / $0.06 / $1.20$0.45 / $0.09 / $1.80Identical card to M2.7
MiniMax M2.7$0.30 / $0.06 / $1.20$0.45 / $0.09 / $1.80Older generation held at the same price
OpenAI GPT OSS 120B$0.15 / $0.015 / $0.60$0.18 / $0.018 / $0.72Smallest Priority uplift on the card (1.2×)
OpenAI GPT OSS 20B$0.07 / $0.035 / $0.30Second-cheapest input rate on the card, behind the new Nemotron 3.5 Lightning row below
Muse Glimmer 30B$0.35 / $0.04 / $1.50$0.525 / $0.06 / $2.25New row as of this capture; Priority is a flat 1.5× uplift on every dimension
NVIDIA Nemotron 3.5 Lightning 30B A3B$0.05 / $0.01 / $0.20New row as of this capture; cheapest published input rate on the entire serverless card, undercutting GPT OSS 20B’s $0.07
NVIDIA Nemotron 3 Ultra (Preview)$0.60 / $0.12 / $2.40Preview model, Standard-only

Serverless inference (other base models, by size and architecture)

Any text or vision model without an individual price row is priced by parameter count and architecture, applied uniformly to input and output with no separate cached-input rate.

Base model classRate per 1M tokensIncludedKey mechanics
Less than 4B parameters$0.10Input and output at the same rateLong-tail hosting for small open models
4B – 16B parameters$0.20Input and output at the same rateDefault band for community fine-tunes
More than 16B parameters$0.90Input and output at the same rateDense-model ceiling
MoE up to 56B (e.g. Mixtral 8x7B)$0.50Input and output at the same rateMoE priced on a separate ladder
MoE 56.1B – 176B (e.g. DBRX, Mixtral 8x22B)$1.20Input and output at the same rateMost expensive size-based band

Batch inference is billed at 50% of serverless pricing on both input and output. Cached input tokens are priced at 50% of the input rate by default for all text and vision language models “unless otherwise specified” — every headline model above specifies a lower cached rate.

On-demand deployments (per GPU hour, billed per GPU second)

Price increase effective September 1, 2026. As of this capture, the on-demand pricing table publishes two columns — the rate through August 31 and the new rate from September 1 — the first repricing of an already-published on-demand GPU tier in the tracked history.

GPU typePrice/hr through Aug 31Price/hr from Sep 1Key mechanics
H100 80 GB$7.00$8.00 (+14%)Single-tenant deployment, autoscaling; no extra charge for start-up time
H200 141 GB$7.00$8.00 (+14%)Same rate as H100 for 1.75× the VRAM; default choice for large-context serving
B200 180 GB$10.00$13.00 (+30%)Blackwell-class capacity; also the substrate for most managed fine-tuning
B300 288 GB$12.00$15.00 (+25%)Largest published VRAM per GPU; frontier-model and longest-context workloads
GB300 288 GB$18.00$20.00 (+11%)Highest published on-demand rate on the card

Default self-serve quota is 8 GPUs each of A100, H100, H200 and B200, plus 100 on-demand LoRAs; larger allocations are a contact-sales request. Reinforcement fine-tuning jobs are billed per GPU hour at these same on-demand rates (at whichever rate is in effect on the billing date). Region-restricted deployments (US, Europe) are priced at a flat 1.5× the standard on-demand rate above and require a Contact Sales request to provision — the first published geographic-routing premium on the dedicated-GPU side of the card, mirroring the existing 10% US-only Serverless premium on the token side; the pricing page does not state whether the 1.5× multiplier applies to the pre- or post-Sep-1 base rate.

Fine-tuning (per 1M training tokens)

Base model sizeLoRA SFTLoRA DPOFull param SFTFull param DPO
Models up to 16B parameters$0.50$1.00$1.00$2.00
Models 16.1B – 80B$3.00$6.00$6.00$12.00
Models 80B – 300B (e.g. Qwen3-235B, GPT-OSS-120B)$6.00$12.00$12.00$24.00
Models >300B (e.g. DeepSeek V3, Kimi K2)$10.00$20.00$20.00$40.00

There is no LoRA surcharge — the pricing page states “Serve fine-tuned models for the same price as base models” — but serving a tuned model is not free and is not serverless. The docs are explicit that fine-tuned LoRA models “can only be deployed to on-demand (dedicated) deployments. Serverless deployment is not supported for LoRA models”, so hosting bills at the per-GPU-second on-demand rate above (from $7.00 per GPU hour): one adapter per live-merge deployment, or up to 100 LoRAs as add-ons on a single multi-LoRA base deployment. Training tokens are estimated as dataset tokens × epochs, and when tuning with intermediate thinking traces that estimate should be multiplied by the average number of conversation turns ÷ 2, because multi-turn conversations are unrolled into user, assistant and thinking traces. Vision (VLM) supervised fine-tuning is also billed per 1M tokens.

Serverless Training API (new 2026-07-29, per 1M tokens)

A Tinker-compatible product, new since the prior capture: attach to a shared, always-on trainer pool for LoRA training on a small set of launch models. No provisioning, no idle cost — billing is purely per token prefilled, sampled, and trained.

Base modelContextPrefill / 1MCached prefill / 1MSample / 1MTrain / 1M
Qwen 3.5 9B64K$0.66$0.132$1.995$1.463
Qwen 3.6 27B128K$1.86$0.372$5.595$4.103
Kimi K3192K$10.87$2.17$27.11$32.55

Checkpoint storage for serverless models is included during private preview; the docs note “other frontier models are coming soon” to the catalog. Dedicated Training API jobs (as opposed to this shared-pool serverless product) are billed per GPU hour at the same on-demand rates listed above.

Embeddings (per 1M input tokens)

Base model parameter countRateIncludedKey mechanics
Up to 150M$0.008Input tokens onlyCheapest retrieval tier
150M – 350M$0.016Input tokens onlyMid-size retrieval models
Qwen3 8B$0.10Input tokens onlyLarge embedding model, priced separately

Fire Pass (experimental)

TierPriceIncludedKey mechanics
Fire PassPromo code activationZero per-token cost on included open-weight models — as of 2026-07-29 the granted model is Kimi K3 Fast, a reasoning model with a 1M-token context window optimized for complex coding tasks, configured as accounts/fireworks/routers/kimi-k3-fast (previously GLM 5.2 Fast at a 256k context window); a dedicated fpk_ API key that only works for those models”Fire Pass is an experimental product. Features, availability, and pricing are subject to change.” Allowed for personal development and agentic coding harnesses (now including Codex, Pi and LangChain Deep Agents alongside Claude Code, OpenCode, Cline, Kilo Code and OpenClaw); prohibited for production workloads, on pain of pass revocation. FireConnect auto-detects fpk_ keys and applies Fire Pass defaults for Claude Code, OpenCode, Codex and Pi

Account spending tiers (legacy postpaid budget ceiling; prepaid sets independently)

TierCriteriaLegacy postpaid max monthly spendTraining GPU quota (Blackwell / H200 / H100+A100)
No payment method0 / 0 / 0
Tier 1Valid payment method and billing profile$500 / 16 / 8
Tier 2Spend or add $50 in credits$50016 / 16 / 16 — first tier with any Blackwell training quota
Tier 3Spend or add $500 in credits$5,00024 / 24 / 24
Tier 4Spend or add $5,000 in credits$50,00032 / 32 / 32
UnlimitedContact usUnlimitedCustom — Enterprise accounts are exempt from tier caps entirely

As of this capture, the docs clarify that the monthly-spend-limit column above applies specifically to legacy self-serve postpaid accounts. Today’s default prepaid accounts instead set their own monthly spend limit independently of tier via firectl quota update monthly-spend-usd; the tier itself still gates serverless TPM upper bounds and training GPU quota regardless of billing mode.

Sales motions across products: PLG / self-serve for serverless, on-demand deployments, fine-tuning and Fire Pass; sales-led for Enterprise accounts, post-paid billing, unlimited spend, custom GPU allocations and region-restricted (US/Europe) on-demand deployments.


Hidden costs : What Fireworks AI customers actually pay beyond the rate card

Archetype A: Developer-tools startup running serverless inference on DeepSeek V4 Flash

A growth-stage AI coding startup running ~90M input + 15M output tokens per month, with high prefix re-use (system prompt + tool definitions cached):

Line itemMonthly cost
Input tokens (90M at the Standard rate, $0.14/1M)$12.60
Cached input (60% prefix hit re-billed at $0.028/1M)-$6.05
Output tokens (15M at $0.28/1M)$4.20
Batch API for offline evaluation runs (50% of serverless)<$5
Estimated total — Standard serving path~$15/month
Same workload on the Priority path ($0.21 / $0.042 / $0.42)~$21/month

Two things drive this bill more than the headline rate. First, the cached rate on this model is $0.028 against a $0.14 input rate — 20% of input, not the 50% platform default — so the cache-hit ratio is the single largest lever a prefix-heavy workload has. Second, since 2026-07-22 the serving-path choice is a real line item: Priority applies a flat 1.5× to every dimension on DeepSeek V4, so buying load-shed protection raises the whole metered bill by half rather than adding a fixed fee. At this scale that is a $6/month decision; at 100× the volume it is the difference between $1,500 and $2,100, which is the point at which it deserves a measurement rather than a default.

Archetype B: Mid-market team running a fine-tuned Llama 70B on dedicated H100

A team that fine-tuned Llama 3.3 70B (full-parameter SFT, 10M training tokens) and serves it on a dedicated H100 with autoscaling:

Line itemMonthly cost
Initial fine-tuning (one-time, 10M tokens × $6.00)$60
H100 dedicated (8h/day × 30 × $7.00)$1,680
Warm-pool retention (avoid cold starts during business hours, ~4h/day)$840
Bandwidth + storage (negligible at this scale)<$10
Estimated total~$2,530/month (after one-time $60 fine-tune)

The H100 dedicated rate dominates the bill — and the customer is paying for warm-pool retention to maintain latency SLAs. The fine-tuning cost is amortized over many months of inference, which makes the unit economics of fine-tuned dedicated deployment far better than per-token serverless once sustained QPS rises above ~10/second.

Want to estimate your own Fireworks AI bill? Use the Fireworks AI pricing calculator to model serverless tokens, dedicated GPU hours, and fine-tuning costs by model size.


Pricing evolution : Fireworks AI’s pricing history from per-token serverless to multi-SKU platform

Cadence

QuarterPrice changesProduct / SKU additionsNotes
2022 Q401Fireworks founded; closed alpha for Llama and Stable Diffusion
2023 Q301Public serverless API launch
2024 Q101Dedicated deployments + LoRA fine-tuning launched
2024 Q300Series B ($52M); no public price changes
2024 Q411Turbo + Priority QoS tiers + cached input 50% discount
2025 Q111Batch API at 50% discount launched
2025 Q301B200 + B300 GPU availability ($10/hr, $12/hr)
2026 Q111Differential embeddings pricing by parameter size
2026 Q347Serverless repriced across three named serving paths; Fast variants and Fire Pass added; Series D + $1B ARR (07-22); Kimi K3 flagship + US-only Serverless premium and a new Serverless Training API added, Fire Pass swapped to Kimi K3 Fast (07-29); GB300 GPU tier + 1.5x region-restricted deployment premium added, first documented exception to the US-only Serverless premium (08-11); on-demand GPU price increase announced, effective Sep 1 (08-12); two new serverless models added, Nemotron 3.5 Lightning now the card’s cheapest input rate (08-14)

Tracked range: 2022 Q4–2026 Q3. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).

Notable changes

  • 2023-08-17 — Public serverless API launch; first major positioning as “faster, cheaper Llama 2 hosting.”
  • 2024-01-31 — Dedicated deployments + LoRA fine-tuning launched; pricing model expanded from single-SKU per-token to multi-SKU platform.
  • 2024-11-04 — Turbo + Priority QoS tiers + cached input 50% discount launched; quality-of-service segmentation became a self-serve API choice rather than a tier upgrade.
  • 2025-03-25 — Batch API launched at 50% discount across all models; compounded with cached input to enable 25%-of-standard pricing on batched cached workloads.
  • 2025-08-19 — B200 + B300 added at $10/hr and $12/hr; Hopper-class H100/H200 remained at $7/hr, positioning Blackwell as a premium for largest-model workloads.
  • 2026-02-10 — Differential embeddings pricing launched; sub-150M-parameter models at $0.008/1M undercut OpenAI text-embedding-3-small by 60%.
  • 2026-07-16 — Series D reported at $1.5B on a $17.5B valuation, earmarked for enterprise inference capacity. In the most price-competitive layer of the AI stack, capital at that scale historically precedes downward per-token pressure rather than list-price increases — and no headline rate moved in the repackaging that followed six days later.
  • 2026-07-22 — Serverless split into three named serving paths. The 2024-era “Turbo + Priority” pairing was replaced by Standard (default), Priority (service_tier parameter, prioritized above Standard traffic and less likely to be load shed) and Fast (separate model ID, targeting 100+ tokens/sec). The substantive change is that each path now carries its own published per-model input / cached-input / output rate instead of an unpriced quality-of-service label: Priority runs 1.2×–1.5× Standard (Kimi K2.6 $0.95 → $1.50 input; GPT OSS 120B $0.15 → $0.18) and Fast runs roughly 2× (Kimi K2.6 Fast $2.00, GLM 5.2 Fast $2.10, GLM 5.1 Fast $2.80, Kimi K2.7 Code Fast $1.90). Reliability and speed became two separately-purchasable axes rather than one bundled premium.
  • 2026-07-22 — Fire Pass launched: promo-code activation, a dedicated fpk_ key, and zero per-token cost on included open-weight models for personal agentic coding — explicitly prohibited for production. It is the first zero-marginal-cost packaging on an otherwise strictly metered platform, and it lands alongside a site banner announcing a Series D and $1B ARR. No headline rate changed: H100/H200 stayed at $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%.
  • 2026-07-29 — Kimi K3 added as a new flagship serverless model at $3.00 / $0.30 / $15.00 (Standard) and $3.75 / $0.375 / $18.75 (Priority), with a Fast variant ($4.50 / $0.45 / $22.50) and a Kimi K3 US variant that is the first model to carry its own priced row for a newly-documented US-only Serverless mechanic: a flat 10% premium over base serverless pricing, described in the docs as applying to any model routed to US-only infrastructure, not just Kimi K3.
  • 2026-07-29 — Fireworks shipped the Serverless Training API, a wholly new product and its first Tinker-compatible offering: LoRA training on a shared, always-on trainer pool, billed per 1M tokens across four dimensions (prefill, cached prefill, sample, train) for three launch models — Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3. Unlike the existing managed fine-tuning grid (priced per training-token by size/method tier as a discrete job) or reinforcement fine-tuning (billed per GPU hour), this product has no provisioning step and no idle cost, metering training the same continuous way Fireworks already meters inference.
  • 2026-07-29 — Fire Pass’s included free model swapped from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context), alongside new supported harnesses (Codex, Pi, LangChain Deep Agents) and FireConnect auto-detection. The per-token price of the pass itself did not change — it is still zero — but the capability behind that zero got materially larger, confirming that Fireworks treats the free grant’s underlying model as a lever it can move independently of the metered rate card.
  • 2026-08-11 — Fireworks added GB300 288 GB to the on-demand GPU card at $18.00/hr, the highest published on-demand rate, and introduced a region-restricted deployments option (US, Europe) priced at a flat 1.5x the standard on-demand rate via Contact Sales — the first geographic-routing premium on the dedicated-GPU side of the product, mirroring the 10% US-only Serverless premium on the token side. The same capture also published the US-only Serverless premium’s first documented exception: GLM 5.2 Fast US is priced identically to global GLM 5.2 Fast ($2.10 / $0.21 / $6.60), with no 10% markup applied — the “flat 10% on any model” framing from 07-29 now needs a per-model check. Separately, the docs clarified that the long-published $50/$500/$5,000/$50,000 spend-tier ceilings apply only to legacy postpaid accounts, not today’s default prepaid accounts.
  • 2026-08-12 — Fireworks published an on-demand GPU price increase effective September 1, 2026, the first repricing of an already-published on-demand SKU since the on-demand product launched in 2024-01. Every GPU tier rises: H100/H200 $7.00 → $8.00/hr (+14%), B200 $10.00 → $13.00/hr (+30%), B300 $12.00 → $15.00/hr (+25%), and GB300 $18.00 → $20.00/hr (+11%). B200 carries the steepest percentage increase on the card. The pricing page frames it as a forward-dated two-column table (current vs. from-Sep-1) rather than an immediate change, giving existing customers roughly three weeks’ notice.
  • 2026-08-14 — Fireworks added two new serverless models: Muse Glimmer 30B ($0.35 / $0.04 / $1.50 Standard, $0.525 / $0.06 / $2.25 Priority — a flat 1.5x uplift) and NVIDIA Nemotron 3.5 Lightning 30B A3B ($0.05 / $0.01 / $0.20, Standard-only). Nemotron 3.5 Lightning’s $0.05 input rate is now the cheapest published on the entire serverless card, undercutting the prior low of OpenAI GPT OSS 20B at $0.07. No other pricing surface changed in this capture — the on-demand GPU increase announced 08-12 is unchanged.

What’s unique : Fireworks AI’s distinctive pricing mechanics

1. Cached input and Batch API discounts stack independently. Most inference platforms offer either cached input OR batch discounts. Fireworks offers both, and they compound: a batched cached workload pays 25% of standard input. For RAG systems and agent loops with high prefix re-use, this is a meaningful structural cost advantage that does not require negotiating a custom contract.

2. Reliability and speed are priced as two separate axes on the same model. Since 2026-07-22 the same weights are sold on three paths with three published rate cards: Standard, Priority (service_tier: "priority" — queue precedence and load-shed protection, 1.2×–1.5× Standard) and Fast (a distinct model ID targeting 100+ tokens/sec, roughly 2× Standard). Almost every competitor bundles “premium” into one faster-and-more-reliable tier; Fireworks makes a buyer who needs 503-resistance during traffic peaks pay only for that, and a buyer who needs interactive token velocity pay only for that. Neither requires a contract or a plan change, which keeps quality-of-service a per-workload usage decision rather than a procurement one.

3. Fire Pass is zero-marginal-cost packaging bolted onto a strictly metered platform. Fireworks has never sold a subscription — every SKU is a meter. Fire Pass (2026-07-22) breaks that pattern in one direction only: a promo code turns off per-token billing entirely on included open-weight models, isolated behind a separate fpk_ key and fenced to non-production personal coding. It is a top-of-funnel instrument disguised as a product, and the key separation is what makes it safe — pass usage and metered usage can never silently mix on one credential. The 2026-07-29 swap of the included model (from GLM 5.2 Fast’s 256k context to Kimi K3 Fast’s 1M context) shows the mechanic is durable rather than a one-off promo: the price of the pass stays fixed at zero while the capability behind it is the lever Fireworks actually moves.

4. Fine-tuning rate card scales linearly with base model size AND training method. Most platforms offer either a flat per-token fine-tuning rate or a per-base-model rate card that ignores training method. Fireworks’ 16-cell grid (4 size tiers × 4 method tiers: LoRA SFT, LoRA DPO, Full SFT, Full DPO) lets cost-conscious teams trade quality for cost with a granularity competitors do not match.

5. Differential embeddings pricing by parameter size. Most embedding APIs (OpenAI text-embedding-3, Voyage, Cohere) price embeddings as a flat rate regardless of model size. Fireworks’ three-tier embeddings schedule ($0.008 / $0.016 / $0.10 per 1M) lets retrieval-pipeline operators pick a model that matches their accuracy target without paying for over-engineered embeddings. For usage-aggregation strategies at scale, the 60–90% savings on small-model embeddings are material.

6. The free entry point moved from credits to a capability grant. The $1 trial credit is still one of the smallest in the market — enough to test that the API responds, not to evaluate a workload. But Fire Pass (2026-07-22) changes what “free” means here: instead of metering a tiny balance down to zero, Fireworks now gives away unmetered access to one strong coding model, and rations it by use case (personal, non-production) rather than by dollars. That is a deliberate pricing-as-positioning choice — it buys daily habit inside Claude Code and Cline where the developer decision actually gets made, while keeping every production token billable.

7. A routing constraint is priced as a flat percentage, not a separate SKU — with one now-documented exception. US-only Serverless (2026-07-29) is the first mechanic on the card that prices where a request runs rather than how fast or how reliably: a model routed to US-only infrastructure costs a flat 10% more than its base serverless price, on every dimension (input, cached input, output) at once — the pattern Kimi K3 US ($3.30 / $0.33 / $16.50 versus Kimi K3’s $3.00 / $0.30 / $15.00) still follows. As of 2026-08-11, though, the docs also published the rule’s first carve-out: GLM 5.2 Fast US is priced identically to global GLM 5.2 Fast ($2.10 / $0.21 / $6.60), with no 10% markup at all. That single exception matters more than its size — it converts “the 10% is a general policy” from an unconditional forecast into a rule a buyer still has to verify per model, which undercuts the same discovery-cost problem this section otherwise credits Fireworks for solving (compare the Priority and Fast multipliers under Areas to improve, which never claimed to be a flat rule in the first place).

8. Training is now metered like inference, not billed like a job. The Serverless Training API (2026-07-29) breaks from Fireworks’ own managed fine-tuning precedent, which prices a discrete job (dataset tokens × epochs, at a flat rate per size/method tier). The new product instead runs a shared, always-on trainer pool and bills per token across four live dimensions — prefill, cached prefill, sample, and train — with no provisioning step and no idle cost, the same continuous-consumption logic Fireworks already applies to serverless inference. It is the clearest sign yet that Fireworks’ pure-usage architecture is a platform-wide default it extends to new resource types, not a policy specific to inference.


Strengths & weaknesses

StrengthsWeaknesses
H100/H200 at $7.00/hr is among the lowest published rates in managed inferenceA100 carries a default self-serve quota of 8 GPUs but has no published per-hour rate
Cached input + Batch API discounts stack (25% of standard for batched cached)$1 trial credit is one of the smallest in the market — limits self-serve evaluation of anything Fire Pass does not cover
Reliability (Priority) and speed (Fast) are separately priced — buy one without the otherPriority uplift is model-specific (1.2× on GPT OSS 120B, 1.5× on Kimi K2.6) with no published rule to forecast against
Every serving path publishes its own input / cached / output rate — no unpriced “premium tier”Fast is a different model ID, not a request parameter — routing between speeds means a code change, not a flag
Fire Pass gives unmetered access to a 256k-context coding model, keyed separately from billable usageFire Pass is “experimental… subject to change” and the non-production boundary is undefined and revocation-backed
Granular fine-tuning grid (4 size × 4 method) for cost-quality trade-offsPer-model serverless rates require docs lookup — not surfaced on the pricing page directly
PyTorch-team founder credibility (Lin Qiao) builds platform-runtime trustImage generation pricing not listed on the on-demand pricing page — only docs
Embeddings pricing tiered by parameter size — 60% cheaper than OpenAI for small modelsNo published path between credit-card serverless and Enterprise commit — friction at scale
US-only Serverless premium is published as a single flat rate (10%) rather than a bespoke per-model surchargeServerless Training API adds a 4th per-token dimension (prefill/cached-prefill/sample/train) on top of the 3-dimension inference card and the 4×4 fine-tuning grid — and the “flat” US-only premium already has a published exception (GLM 5.2 Fast US carries no markup), so a buyer sizing Kimi K3 end-to-end still reconciles well over a dozen numbers with at least one per-model check
Serverless Training API has no provisioning step or idle cost — training bills only for tokens actually processedTraining API launched with only 3 models (Qwen 3.5 9B, Qwen 3.6 27B, Kimi K3) and is in private preview — most fine-tuning workloads still route to the older, coarser managed fine-tuning grid

Billing UX : Fireworks AI’s account controls and payment experience

  • firectl quota list — The canonical read command: prints the account’s rate limits, GPU quotas, spend limits and usage across serverless and on-demand deployments in one place.
  • firectl quota update monthly-spend-usd --value <AMOUNT> — Self-serve monthly budget cap, adjustable at any time (the docs example sets a $200 monthly budget). The cap cannot exceed the account’s spending-tier ceiling.
  • Automatic spend pause — “When you reach your spending limit, all API requests pause automatically” across serverless inference, deployments and fine-tuning. Resuming requires adding credits or raising the budget cap. Explicitly does not apply to Enterprise accounts, which are never paused for spend.
  • Spending tiers as a self-service unlock — Adding prepaid credits promotes the account (adding $100 moves Tier 1 → Tier 2) and the new tier “activates within minutes”, simultaneously lifting the serverless TPM upper bounds and the training-GPU allocation. As of this capture, the docs now distinguish the tier’s budget-ceiling effect from its capacity effect: the published $50/$500/$5,000/$50,000 monthly-spend ceilings apply only to legacy self-serve postpaid accounts, while today’s default prepaid accounts set their own spend limit independently of tier.
  • Auto Reload — A separate control from the monthly spend limit: it purchases credits automatically when the account balance runs low. It does not raise or reset the monthly spend limit itself.
  • Account-wide request ceiling — 10 RPM with no payment method or no credits; 6,000 RPM once a payment method and active credits exist. The 6,000 RPM cap is fixed rather than adaptive and is shared across all API usage, so per-minute volume above it is rejected with HTTP 429 regardless of spending tier.
  • Training GPU quota, granted by tier — A separate pool from on-demand deployment quota, now published per GPU family: 0 Blackwell / 0 H200 / 0 H100+A100 with no payment method; 0 / 16 / 8 at Tier 1; 16 / 16 / 16 at Tier 2 (first Blackwell access); 24 across the board at Tier 3; 32 across the board at Tier 4 — with rejected jobs returning HTTP 429 quota_exceeded. Because current managed fine-tuning shapes run on Blackwell, most fine-tuning requires reaching Tier 2.
  • Prepaid credits by default — “Fireworks operates on a pre-paid credits billing system”; moving to post-paid is a contracted-customer option negotiated with sales.
  • Exporting Billing Metrics + Usage & Cost Breakdown — Two dedicated Administration doc surfaces, backed by firectl billing export-metrics, firectl billing get-usage, firectl billing list-invoices and firectl billing notification-settings for programmatic usage, rated cost, invoice and alert management.
  • Per-User Usage Limits (Fireworks for Work) — Per-user spending limits on serverless inference with account defaults and per-user overrides, for teams sharing one Fireworks account.
  • Fire Pass key separation — Fire Pass issues a distinct fpk_... API key that only works for included open-weight models; the normal Fireworks key continues to bill all other models, so pass usage and metered usage never mix on one credential.
  • Automatic account recovery — Suspended accounts reactivate automatically within an hour once failed payment methods are resolved and credits added (or outstanding invoices paid, for post-paid accounts).

Strategic wins : Why Fireworks AI’s pricing decisions worked

1. PyTorch-team founder credibility as the runtime-optimization moat

Lin Qiao led the PyTorch team at Meta and shipped PyTorch 1.0; her co-founding of Fireworks gave the company immediate credibility with ML engineering teams evaluating inference runtimes. When customers compare Fireworks’ FireAttention kernels and FireOptimizer auto-tuning claims against open-source vLLM or TensorRT-LLM, the founder credentials act as a trust-multiplier that pure-marketing positioning cannot replicate. This made the $7/hr H100 rate believable rather than skeptical.

2. Stackable discounts (Cached input + Batch) as a competitive moat

By making cached input and Batch API discounts compound independently, Fireworks created a structural pricing advantage for RAG and agent-loop workloads that competitors offering only one or the other cannot match. For sustained high-volume customers with prefix re-use, the 25%-of-standard effective rate makes Fireworks materially cheaper at scale — a pricing-mechanic moat that procurement leaders can validate without sales conversations.

3. Unbundling reliability from speed turned quality-of-service into two priced meters

The 2024 Turbo + Priority tiers already proved that customers would self-select quality-of-service without a contract. The 2026-07-22 rebuild finished the job: Standard, Priority and Fast each publish their own per-model input / cached-input / output rate, so the buyer no longer picks a labelled tier — they price two independent questions. Do I need to survive a traffic peak without 503s? costs 1.2×–1.5×. Do I need 100+ tokens/sec? costs about 2×. Because Priority is a service_tier parameter and Fast is a model ID, both remain per-request choices inside one account rather than seat-tier or commitment structures — a team can route its interactive path to Fast and leave its background jobs on Standard without negotiating anything.

4. Granular fine-tuning grid captured cost-conscious teams competitors missed

The 16-cell fine-tuning rate card (4 size tiers × 4 method tiers) gave cost-conscious teams a way to dial cost versus quality with precision: pick LoRA for cheap iteration, full-parameter for production quality, SFT for instruction following, DPO for preference alignment. Most competitors’ coarse-grained fine-tuning pricing forced over-payment for unnecessary capabilities. The granular grid let Fireworks capture both “fast cheap iteration” and “high-quality production” buyers with the same SKU surface.

5. A pure-usage rate card reached $1B ARR with no subscription line

The July 2026 banner claims $1B ARR alongside a Series D (reported 2026-07-16 at $1.5B on a $17.5B valuation). Fireworks sells no seats, no platform fee and no minimum — the entire run-rate is metered consumption drawn against prepaid credits, across four meters (tokens, GPU-seconds, training tokens, embedding tokens). That is a direct counter to the standing objection that pure-usage pricing cannot produce predictable, financeable revenue at scale. The mechanism is visible in the rate card itself: spending tiers cap the vendor’s credit risk, prepaid credits collect cash before consumption, and Enterprise accounts exit both once they are creditworthy — so revenue quality improves as accounts grow rather than degrading with usage volatility.

6. Protocol compatibility as a distribution wedge, reused from inference to training

Fireworks built its inference business partly by matching the OpenAI- and Anthropic-compatible API shapes, so switching a workload over meant changing a base URL, not rewriting a client. The 2026-07-29 Serverless Training API repeats that exact playbook one layer up the stack: it is Tinker-compatible, meaning teams already writing training loops against that interface can point them at Fireworks’ shared trainer pool with minimal integration work. Pairing that with per-token billing (prefill/cached-prefill/sample/train) rather than a provisioned-GPU rental turns training evaluation into the same zero-commitment motion that made serverless inference self-serve — a customer can test a LoRA run without booking a GPU-hour block first.


Areas to improve : Gaps in Fireworks AI’s pricing approach

1. Per-model serverless rates should be on the pricing page, not buried in docs

The pricing page lists discount mechanics (cached, batch) and explains the serving paths, but the actual per-model rates live on the docs serverless-pricing card. The 2026-07-22 repackaging made this worse rather than better: a buyer now needs three numbers per model per path — input, cached input and output, across Standard, Priority and Fast — which is up to nine figures before they can size a workload, and none of them are on the page they land on. Surfacing the top-10-model card inline would materially improve self-serve conversion.

2. Priority and Fast multipliers are per-model, with no published rule

The uplift is not a constant. Priority is 1.2× on GPT OSS 120B ($0.15 → $0.18), 1.25× on GLM 5.2 ($1.40 → $1.75), 1.5× on Kimi K2.7 Code and DeepSeek V4 — and on Kimi K2.6 the input uplift is 1.58× while the cached uplift is only 1.375×, so the effective premium moves with a workload’s cache-hit rate. Fast is roughly 2× but GLM 5.1 Fast is exactly 2× on input while Kimi K2.6 Fast is 2.1×. A team cannot budget “Priority costs 1.5× more” and be right; every model has to be re-priced by hand. Publishing the multiplier as a stated policy — or holding it constant per path — would make reliability spend forecastable instead of a per-model lookup.

3. Fire Pass needs a defined boundary, not just a prohibition

Fire Pass carries real terms — non-production only, violations “may result in pass revocation” — but publishes no usage ceiling, no definition of what makes a workload “production”, and an explicit warning that features, availability and pricing are all subject to change. That is a difficult foundation for a developer deciding whether to build a habit on it, and a difficult one for Fireworks to enforce consistently. A published rate limit or monthly token allowance would convert an ambiguous prohibition into a boundary both sides can see, without giving up the zero-marginal-cost hook.

4. $1 trial credit is too small for meaningful evaluation

Together AI typically provides $5 trial credit; Anyscale offers $100; Baseten provides variable credits that historically reached $30. Fireworks’ $1 trial covers a few thousand tokens of inference but is insufficient to evaluate even a single non-trivial RAG workflow. Increasing to $5–$25 would let self-serve developers reach proof-of-value without sales conversations and likely accelerate self-serve revenue growth.

5. A100 pricing not on the on-demand page creates a coverage gap

The on-demand dedicated page lists only H100/H200, B200, and B300 — no A100. Yet the account-quotas doc grants every self-serve account a default quota of 8 A100 GPUs, so buyers can provision hardware whose price they cannot look up and must ask sales for. Publishing an A100 rate (even at a “limited availability” disclaimer) would broaden the on-demand TAM and reduce friction for non-frontier workloads.

6. No published bridge between credit-card self-serve and Enterprise commit

Self-serve customers grow into the $10K–$50K/month band before hitting the Enterprise commit tier — and there is no published mid-tier discount schedule for this band. Publishing a volume discount ladder (e.g., 10% off above $10K/mo, 15% above $25K/mo) would let mid-market customers self-qualify without sales friction and reduce churn at the boundary.

7. Each new product launch adds a full row of mechanics buyers must independently discover

The 2026-07-29 additions compound rather than simplify the rate card: Kimi K3 shipped with three priced variants (Standard, Fast, US), the Serverless Training API introduced a fourth per-token dimension nowhere else on the card (prefill, cached prefill, sample, train), and the US-only Serverless premium is documented on its own separate page rather than inline with the model it first prices. None of this is on the pricing page itself — a buyer sizing a Kimi K3 workload across inference, US routing, and training now needs to cross-reference at least three separate docs pages to assemble one number. A single “what’s new” or “total cost” summary at the top of the serverless pricing page would let the granularity keep paying off without the discovery cost rising every time Fireworks ships.


Monetization stack & signals : how Fireworks AI builds & buys its revenue engine

Buys 5 Builds 1 6 open roles

The read — where the monetization investment is going

Fireworks builds the usage→revenue spine in-house — a Data Platform Engineer owns the order-to-cash pipeline on a BigQuery + dbt warehouse — but buys the pieces around it: Orb for billing, NetSuite as ERP/GL, Salesforce as CRM. The revenue org is staffing RevOps, ASC-606 revenue accounting, and growth.

Stack — build vs buy
Builds in-house · 1
  • In-house usage metering & order-to-cash pipeline In-house build inferred Job post 1 Docs 2 Jun 2026

    “...build and own the end-to-end OTC data pipeline (usage → billing → payments → revenue → GL → reporting) ... usage metering and invoice generation through revenue recognition.”

Buys (vendor) · 5
  • Salesforce CRM Job post 1 Job post 2 Job post 3 Jun 2026

    “Proficiency in CRM and marketing automation software; Salesforce is required.”

  • BigQuery Data platform Job post Jun 2026

    “Strong SQL and BigQuery proficiency — you can design schemas, write complex analytical queries, build dbt models, and maintain production data pipelines.”

  • dbt Data platform Job post Jun 2026

    “...write complex analytical queries, build dbt models, and maintain production data pipelines.”

  • Orb Billing inferred Job post 1 Job post 2 Jun 2026

    “You will work hands-on with our billing platform (Orb, etc).”

  • NetSuite Revenue recognition inferred Job post 1 Job post 2 Jun 2026

    “Working knowledge of accounting systems (QuickBooks or NetSuite) and the ability to map billing events to GL journal entries, manage sub-ledger reconciliation, and support month-end close.”

Open roles in the revenue & lifecycle org — 6
View open roles
  • Member of Technical Staff, Data Platform Engineer Billing engineeringData platform seen Jun 11, 2026
  • Director, Revenue Strategy & Analytics Data platformRevOpsRetention seen Jun 11, 2026
  • Head of Marketing Operations RevOpsGrowth seen Jun 11, 2026
  • Sales Strategy Lead RevOps seen Jun 11, 2026
  • Paid Growth Marketer Growth seen Jun 11, 2026
  • Revenue Accounting Lead Deal desk seen Jun 4, 2026

Signals reviewed · derived from public job posts, product docs

Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.

Key takeaways

  1. Founder credibility in the runtime layer is the most durable inference-platform moat. Lin Qiao’s PyTorch leadership made Fireworks’ optimization claims believable in a way that pure-marketing positioning cannot replicate. Infrastructure commercializations that lack maintainer or core-team alignment must invest disproportionately in published benchmarks to compensate.

  2. Stackable discounts (Cached + Batch) are a structural competitive moat for RAG. Independent stacking of cached input and batch discounts lets Fireworks land at 25% of standard for prefix-heavy workloads — a cost-advantage durability that competitors offering only one discount cannot match without contract negotiation.

  3. Price reliability and speed separately — they are not the same premium. Fireworks’ 2026-07-22 split into Standard / Priority / Fast shows that “premium tier” was hiding two unrelated purchases: load-shed protection during peaks (1.2×–1.5×) and token velocity (~2×). Bundling them makes every buyer overpay for the half they do not need. The caveat is enforcement mechanics: Priority is a request parameter and Fast is a separate model ID, so only one of the two can be toggled without a code change.

  4. Granular fine-tuning grids capture both cost-conscious and quality-conscious buyers. The 4×4 size × method matrix lets a single rate card serve fast-iteration teams (LoRA SFT) and production-quality teams (full DPO) without forcing oversized SKU choices. The 2026-07-29 Serverless Training API extends the same granularity philosophy further: instead of a job-based rate per size/method tier, it meters LoRA training continuously per token (prefill, cached prefill, sample, train) with no idle cost — proof that granular value-metric pricing generalizes from inference to training, not just within a single fine-tuning SKU.

  5. Free access can be rationed by use case instead of by dollars. Fireworks pairs a $1 trial credit with Fire Pass, which is unmetered but fenced to personal, non-production coding. Rationing by permitted use rather than by balance means the free tier can be generous exactly where habit forms (a developer’s coding harness) while every production token stays billable — a cleaner segmentation than the trial-credit dial, provided the boundary is enforceable.


UBP implications

  1. Discount stacking is the next frontier in token-based pricing competition. When cached input becomes table stakes, the marginal differentiator shifts to which platforms stack additional discounts (batch, time-of-day, commitment). Usage-based platforms should design discount mechanics to compound independently rather than override each other.

  2. Every dimension of quality-of-service or resource type can carry its own meter. Fireworks’ three serving paths each publish input / cached / output rates, which turns “do I need Priority?” into arithmetic instead of a sales conversation — and the 2026-07-29 Serverless Training API applies the identical idea to training, metering prefill, cached-prefill, sample, and train tokens continuously instead of billing per job. Any usage-aggregated billing surface running heterogeneous workloads or resource types on one key can copy this — but the lesson from the per-model multipliers (1.2×–1.5× on Priority, ~2× on Fast) is that the uplift must be stated as a rule, not left to be reverse-engineered per SKU, or buyers cannot forecast it.

  3. A zero-marginal-cost pass is a viable top-of-funnel for a fully metered product. Fire Pass suspends billing on one model, behind its own key, for one declared use case — the first non-metered thing Fireworks has ever sold. The design generalizes to any pure-usage vendor whose trial credit is too small to demonstrate value: isolate the giveaway on a separate credential so it can never contaminate metered usage, scope it by permitted use rather than by credit balance, and label it experimental so the pricing stays reversible. The open question every vendor copying it must answer is what happens when the free habit meets a production deployment.


Sources


Bottom line

Fireworks AI built its pricing architecture on three structural ideas: that PyTorch-team founder credibility justifies aggressive per-hour GPU pricing ($7/hr H100 is among the lowest in managed inference), that stackable discounts (cached input + Batch API) create a structural cost moat for RAG and agent-loop workloads, and that quality-of-service belongs in the rate card rather than in a contract. The 2026-07-22 rebuild pushed that third idea further than anyone else in the category: Standard, Priority and Fast now each publish their own per-model input / cached / output rates, so reliability and speed are bought separately. The granular fine-tuning grid (4×4 size × method) and differential embeddings pricing (by parameter count) extend the same granularity philosophy to adjacent SKUs. One week later, on 2026-07-29, Fireworks applied the identical philosophy to two dimensions it had never priced before: geography, via a flat 10% US-only Serverless premium that applies broadly rather than a one-off Kimi K3 surcharge (GLM 5.2 Fast US is the one published exception, carrying no markup), and training, via a Serverless Training API that meters LoRA training per token with no provisioning or idle cost instead of billing it as a discrete job.

For AI engineering teams running prefix-heavy RAG workloads or evaluating between cost and quality on fine-tuning, Fireworks delivers one of the most legible commercial inference platforms in the market — and the July 2026 claim of $1B ARR with no subscription line anywhere in the model is the clearest proof in the corpus that a pure-usage rate card can carry a business at scale. The cost of that granularity is arithmetic: nine published figures per model across three paths, with Priority multipliers that vary by model (1.2×–1.5×) and no stated rule to forecast against — and as of 2026-07-29 a buyer evaluating Kimi K3 end-to-end (Standard, Fast, US, plus the new Training API) now reconciles well over a dozen figures across three separate docs pages. The remaining gaps ($1 trial too small, per-model rates buried in docs, no published mid-market discount ladder, A100 missing from on-demand, Fire Pass’s non-production boundary undefined) are GTM polish problems rather than structural pricing flaws.

Compare with peers via the blueprint corpus, or model your own spend with the Fireworks AI pricing calculator.

Pricing timeline : Major events on a vertical axis

Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.

DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added

Fireworks silently repriced DeepSeek V4 Flash (0731) on the serverless rate card — Standard input rose from $0.14 to $0.22 per 1M tokens (+57%), cached input fell from $0.028 to $0.007 (−75%), and output rose from $0.28 to $0.66 (+136%); Priority moved from $0.21 / $0.042 / $0.42 to $0.275 / $0.00875 / $0.825 on the same dimensions. The separate undated "DeepSeek V4 Flash" row (previously identical to the (0731) row) has been removed, leaving the dated row as the sole DeepSeek V4 Flash entry. Unlike the forward-dated, announced on-demand GPU increase from 2026-08-12, this repricing carries no effective-date framing on the page — it reads as already in effect. Separately, a new row, Qwen 3.8 Max, was added to the serverless card at $2.00 / $0.25 / $6.00 (Standard-only, no Priority path).

DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added - Fireworks silently repriced DeepSeek V4 Flash (0731) on the serverless rate card
captured

Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning

Fireworks quietly added two new rows to its serverless per-model rate card: Muse Glimmer 30B ($0.35 / $0.04 / $1.50 Standard, $0.525 / $0.06 / $2.25 Priority — a flat 1.5x uplift) and NVIDIA Nemotron 3.5 Lightning 30B A3B ($0.05 / $0.01 / $0.20, Standard-only). Nemotron 3.5 Lightning's $0.05 input rate is now the cheapest published on the entire serverless card, undercutting the previous low of OpenAI GPT OSS 20B at $0.07. No other pricing page or docs surface changed in this capture.

Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning - Fireworks quietly added two new rows to its serverless per-model rate card: Muse
captured

On-demand GPU price increase announced, effective September 1, 2026

Fireworks published a forward-dated price increase on every on-demand GPU tier: H100/H200 from $7.00 to $8.00/hr (+14%), B200 from $10.00 to $13.00/hr (+30%), B300 from $12.00 to $15.00/hr (+25%), and GB300 from $18.00 to $20.00/hr (+11%), effective September 1, 2026. It is the first repricing of an already-published on-demand SKU since the product launched in January 2024 — every prior on-demand pricing change had been a new-tier addition (A100/H100 in 2024, B200/B300 in 2025, GB300 in early August 2026), not a rate change on an existing GPU. The pricing page shows the current and future rates side by side in a two-column table.

On-demand GPU price increase announced, effective September 1, 2026 - Fireworks published a forward-dated price increase on every on-demand GPU tier:
captured

GB300 GPU tier + region-restricted deployment premium added

Fireworks added GB300 288 GB to the on-demand dedicated GPU card at $18.00/hr, above B300's $12.00/hr — the highest published on-demand rate on the card. The on-demand pricing page also now documents a region-restricted deployments option (US, Europe) priced at a flat 1.5x the standard on-demand rate, requiring a Contact Sales request — the first geographic-routing premium on the dedicated-GPU side of the product, mirroring the existing 10% US-only Serverless premium on the token side. Separately, the docs now note GLM 5.2 Fast US is exempt from that 10% token-side premium (priced identically to global GLM 5.2 Fast), and clarify that the long-published $50/$500/$5,000/$50,000 spend-tier ceilings apply only to legacy self-serve postpaid accounts, not today's default prepaid accounts.

GB300 GPU tier + region-restricted deployment premium added - Fireworks added GB300 288 GB to the on-demand dedicated GPU card at $18.00/hr, a
captured

Kimi K3 flagship, US-only Serverless premium, and Serverless Training API

Fireworks added Kimi K3 (plus a Fast variant and a US-only variant) to its serverless rate card, publishing the first priced row for a new US-only Serverless mechanic — a flat 10% premium over base serverless pricing that the docs describe as applying to any model. Fireworks also shipped the Serverless Training API, a Tinker-compatible product that meters LoRA training per token (prefill, cached prefill, sample, train) on a shared trainer pool with no provisioning or idle cost, covering Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3. Fire Pass's included free model swapped from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context).

Kimi K3 flagship, US-only Serverless premium, and Serverless Training API - Fireworks added Kimi K3 (plus a Fast variant and a US-only variant) to its serve
captured

Three serving paths + Fire Pass; Series D and $1B ARR

Serverless inference is now documented as three named serving paths — Standard (default), Priority (higher reliability under peak traffic, set via service_tier) and Fast (100+ tokens/sec, selected by model ID) — replacing the earlier "Turbo + Priority" framing, with per-model input / cached-input / output rates published for each. Fireworks also shipped Fire Pass, an experimental promo-code pass that removes per-token charges on included open-weight models for personal agentic coding, and the site banner now announces a Series D and $1B ARR. Headline rate card (H100/H200 $7.00/hr, B200 $10.00/hr, B300 $12.00/hr, fine-tuning from $0.50 per 1M training tokens, batch at 50%) is unchanged.

Three serving paths + Fire Pass; Series D and $1B ARR - Serverless inference is now documented as three named serving paths — Standard (
captured

Embeddings Pricing by Parameter Size

Fireworks published differential embeddings pricing by base model parameter count: <150M params at $0.008/1M tokens, 150–350M at $0.016/1M, Qwen3 8B at $0.10/1M. The schedule undercuts OpenAI text-embedding-3-small ($0.02/1M) by 60% for the smallest tier and creates a granular cost ladder for retrieval-pipeline cost optimization.

Embeddings Pricing by Parameter Size screenshot 1
Embeddings Pricing by Parameter Size screenshot 2

B200 and B300 GPU Availability — Frontier Pricing

Fireworks added NVIDIA B200 (180GB) at $10.00/hr and B300 (288GB) at $12.00/hr on-demand dedicated. The H100/H200 rate remained at $7.00/hr, positioning B200/B300 as a premium for largest-model workloads while keeping Hopper-class pricing as the volume default.

Batch API at 50% Discount Across All Models

Fireworks launched a Batch API for asynchronous inference at a flat 50% discount versus serverless. Combined with cached input, batched RAG workloads can land at 25% of standard rates. The Batch API mirrors OpenAI's batch pricing structure and competes directly with Together AI's batch tier.

Turbo + Priority Tiers + Cached Input Discount

Fireworks introduced Turbo (latency-optimized) and Priority (throughput-optimized) tiers on serverless inference, with cached input tokens discounted 50% on both. The tiers let customers self-select for low-latency interactive workloads versus high-throughput batch workloads from the same API.

Series B ($52M) at $552M Valuation

Fireworks raised a $52M Series B led by Sequoia Capital at a $552M post-money valuation. Bessemer, Benchmark, and NVIDIA participated. The round funded the launch of Turbo + Priority quality-of-service tiers and the FireOptimizer auto-tuning service.

Dedicated Deployments + LoRA Fine-Tuning

Fireworks added on-demand dedicated GPU deployments (per-hour A100, H100) and LoRA-based fine-tuning. Pricing established the per-1M-training-token model that remains the canonical fine-tuning SKU, with rates scaling by base model parameter count.

Public Launch — Per-Token Serverless API

Fireworks launched its serverless API publicly, offering Llama 2, Code Llama, Mistral, and Stable Diffusion at competitive per-token rates. Pricing was a flat per-million-token rate that varied by model. Positioned as a faster, lower-cost alternative to OpenAI for open-source workloads.

Fireworks AI Founded

Lin Qiao (former PyTorch team lead at Meta), Dmytro Ivchenko, and Pawel Garbacki founded Fireworks AI to build a high-performance serving platform for open-source generative models. Initial product was a developer preview of optimized inference for Llama 2 and Stable Diffusion.

Trivia
  • · Fireworks AI's $7.00/hour H100 (and H200) on-demand price is one of the lowest published rates among managed inference platforms — roughly 30% below Together AI's $5.49–$6.49 H100 dedicated rates only because Together's listed rate excludes Fireworks' full-stack optimization layer.
  • · Fireworks was founded in 2022 by Lin Qiao (ex-Meta), Dmytro Ivchenko, and Pawel Garbacki — Lin Qiao led the PyTorch team at Meta when PyTorch 1.0 shipped, giving Fireworks unusual inference-runtime credibility.
  • · Fireworks' fine-tuning rate card is one of the most granular in the industry: LoRA SFT at $0.50 per 1M training tokens for <16B models, LoRA DPO at $1.00, full-parameter SFT at $1.00, full-parameter DPO at $2.00 — and it scales linearly through the 16B → 300B+ model size tiers.

Questions & answers

How much does Fireworks AI cost per month?
Fireworks has no monthly subscription fee — you pay only for serverless tokens consumed, dedicated GPU hours running, and fine-tuning training tokens processed. A small RAG application on DeepSeek V4 Flash (0731) ($0.22/1M input, $0.66/1M output) at 30M input + 10M output tokens costs roughly $13/month on serverless; the same workload on a dedicated H100 ($7.00/hr) running 4h/day would cost about $840/month.
What does Fireworks charge per GPU hour for dedicated deployments?
Through August 31, 2026, Fireworks publishes four on-demand dedicated rates: H100 80GB / H200 141GB at $7.00/hour, B200 180GB at $10.00/hour, B300 288GB at $12.00/hour, and GB300 288GB at $18.00/hour. As of this capture, the pricing page also publishes a price increase effective September 1, 2026: H100/H200 rise to $8.00/hour (+14%), B200 to $13.00/hour (+30%), B300 to $15.00/hour (+25%), and GB300 to $20.00/hour (+11%) — the first repricing of an existing on-demand SKU Fireworks has published. There is no published A100 rate on the on-demand page, even though the docs grant every self-serve account a default quota of 8 A100 GPUs — the A100 rate has to be quoted. Dedicated deployments include the platform optimization layer (request batching, KV cache management) on top of the GPU rate.
How do Fireworks Standard, Priority and Fast serving paths differ?
Standard is the default path and needs no extra parameter. Priority is for workloads that need higher reliability during peak traffic — it is prioritized above Standard traffic and less likely to be load shed, and costs more (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority). Fast is a higher-speed path targeting 100+ tokens/sec on the same model at roughly double the Standard rate (Kimi K2.6 Fast is $2.00/1M input). Priority is set with the service_tier parameter; Fast is selected by using a different model ID. As of 2026-07-29 a fourth mechanic, US-only Serverless, prices where a request runs rather than how fast or reliably: routing to US-only infrastructure costs a flat 10% premium over the base model's serverless price on most models (Kimi K3 US is $3.30/1M input versus Kimi K3's $3.00/1M) — though Fireworks has already published one exception, as of 2026-08-11: GLM 5.2 Fast US carries no premium and matches the global GLM 5.2 Fast rate exactly.
Does Fireworks have a free tier, and is Fire Pass really free?
There are two different free entry points. New accounts receive $1 in trial credits — enough to confirm the API works, not enough to evaluate a real workload. Fire Pass is the other one, and it genuinely costs nothing per token on the open-weight models it includes; it is activated with a promo code rather than a payment method. The constraints are the price: Fire Pass is limited to personal development and agentic coding harnesses and is explicitly prohibited for production workloads, with violations able to result in the pass being revoked, and it runs on a dedicated fpk_ API key so every other model still bills against your normal Fireworks key. Fireworks also labels it experimental, with features, availability and pricing subject to change — so it is a good way to build a habit, not a foundation for a product.
How does Fireworks fine-tuning pricing work?
Fine-tuning is priced per 1M training tokens, with rates scaling by base model parameter count and training method. For <16B models: LoRA SFT $0.50, LoRA DPO $1.00, full SFT $1.00, full DPO $2.00. For 16–80B models the schedule scales to $3–$12; 80–300B at $6–$24; >300B at $10–$40. Serving the tuned model carries no LoRA surcharge, but it is not serverless: Fireworks documents that fine-tuned LoRA models can only be deployed to on-demand (dedicated) deployments, so hosting bills at the per-GPU-second on-demand rate (from $7.00 per GPU hour) on top of the one-time training cost. A separate product, the Serverless Training API (added 2026-07-29), instead meters LoRA training continuously per token — prefill, cached prefill, sample, and train — on a shared, always-on trainer pool with no provisioning or idle cost, currently covering Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3.
What discounts apply to batch and cached inference?
Batch inference is billed at 50% of serverless pricing on both input and output. Cached input tokens are priced at 50% of the input rate by default for all text and vision language models "unless otherwise specified" — and the headline models on the docs price card are specified far lower (Kimi K2.6 caches at $0.16/1M against a $0.95/1M input rate, DeepSeek V4 Pro at $0.145 against $1.74). The two discounts stack, so a batched prefix-heavy workload can land well under half the standard rate.