Fireworks AI Pricing Calculator
Updated August 2026Pure-usage per-token serverless + per-hour GPU + per-1M-token fine-tuning + sales-led enterprise
Fireworks AI pricing: Fireworks AI pricing 2026: serverless per token (Standard/Priority/Fast), on-demand GPUs $7-18/hr rising to $8-20/hr Sep 1, fine-tuning per 1M tokens, batch 50% off. Use the free calculator below to enter your usage and get an instant all-in monthly estimate — including overages — and see which plan is cheapest for your needs.
How Fireworks AI prices: Pure-usage per-token serverless + per-hour GPU + per-1M-token fine-tuning + sales-led enterprise.
How Fireworks AI prices
Pure-usage per-token serverless + per-hour GPU + per-1M-token fine-tuning + sales-led enterprise
Your usage
💡 1M input tokens ≈ 750k words of prompts/context; a busy RAG app runs 50–200M/mo
For models without an individual price row, Fireworks sets serverless price by base-model parameter count and applies it uniformly to input and output (no separate cached-input rate). Batch inference bills at 50% of serverless.
💡 Output tokens are the model's generated response; typically 20–40% of input volume for RAG
Per-model output rates from the Fireworks docs serverless price card. Output tokens always cost more than input; match this to the same model class as your input rate.
💡 One GPU running 8h/day ≈ 240 GPU-hrs/mo; 24×7 single GPU ≈ 730 GPU-hrs/mo
On-demand dedicated deployments are billed per GPU-second (shown per hour). No charge for start-up time; idle running time is billed while the deployment is up. Region-restricted deployments (US, Europe) price at 1.5x these rates and require Contact Sales. Rates shown are current through Aug 31, 2026 — Fireworks has published a price increase effective Sep 1, 2026 (H100/H200 to $8.00, B200 to $13.00, B300 to $15.00, GB300 to $20.00).
💡 Training tokens ≈ dataset tokens × epochs; a 10M-token full-SFT run on a 70B model costs ~$60
Fireworks fine-tuning is priced per 1M training tokens by base-model size and method (LoRA vs full-parameter, SFT vs DPO). Serving the tuned model bills separately: LoRA models can only run on an on-demand dedicated deployment (from $7.00 per GPU hour), not on serverless.
Monthly Estimate
Pay-as-you-go
$36.00
per month
Cost Breakdown
Annual
$432.00
Why Pay-as-you-go — $36.00/mo
- Pay-as-you-go at $36.00/mo.
- How we got there: 90 M tokens/mo (Serverless input tokens) × 0 $/1M in (Serverless model size (input rate)) = 18 $.
- How we got there: 30 M tokens/mo (Serverless output tokens) × 1 $/1M out (Serverless model (output rate)) = 18 $.
- How we got there: 0 GPU-hrs/mo (Dedicated GPU hours) × 7 $/hr (GPU type (per-hour rate)) = 0 $.
- How we got there: 0 M tokens (Fine-tuning training tokens) × 1 $/1M (Fine-tuning size × method (per 1M tokens)) = 0 $.
- Serverless input tokens (18 $ × $1) is 50% of your bill — the lever that matters most.
- Need less? Serverless model size (input rate): Under 4B params — $0.10/1M drops you to Pay-as-you-go at $27.00/mo.
Budget range
Plan for usage swings, not just today's estimate.
Conservative
$27.00
45 M tokens/mo/mo
Expected
$36.00
90 M tokens/mo/mo
Aggressive
$54.00
180 M tokens/mo/mo
Need a calculator like this on your pricing page?
Embed interactive pricing calculators on your website to help customers understand costs and boost conversions.
Get StartedAbout this Fireworks AI calculator
This calculator estimates your Fireworks AI cost from publicly available pricing. Actual costs may vary with your specific agreement, volume discounts, and usage patterns — always verify on the provider's official pricing page for the most current rates.
Fireworks AI pricing — frequently asked questions
How much does Fireworks AI cost per month? ▼
Fireworks has no monthly subscription fee — you pay only for serverless tokens consumed, dedicated GPU hours running, and fine-tuning training tokens processed. A small RAG application on DeepSeek V4 Flash (0731) ($0.22/1M input, $0.66/1M output) at 30M input + 10M output tokens costs roughly $13/month on serverless; the same workload on a dedicated H100 ($7.00/hr) running 4h/day would cost about $840/month.
What does Fireworks charge per GPU hour for dedicated deployments? ▼
Through August 31, 2026, Fireworks publishes four on-demand dedicated rates: H100 80GB / H200 141GB at $7.00/hour, B200 180GB at $10.00/hour, B300 288GB at $12.00/hour, and GB300 288GB at $18.00/hour. As of this capture, the pricing page also publishes a price increase effective September 1, 2026: H100/H200 rise to $8.00/hour (+14%), B200 to $13.00/hour (+30%), B300 to $15.00/hour (+25%), and GB300 to $20.00/hour (+11%) — the first repricing of an existing on-demand SKU Fireworks has published. There is no published A100 rate on the on-demand page, even though the docs grant every self-serve account a default quota of 8 A100 GPUs — the A100 rate has to be quoted. Dedicated deployments include the platform optimization layer (request batching, KV cache management) on top of the GPU rate.
How do Fireworks Standard, Priority and Fast serving paths differ? ▼
Standard is the default path and needs no extra parameter. Priority is for workloads that need higher reliability during peak traffic — it is prioritized above Standard traffic and less likely to be load shed, and costs more (Kimi K2.6 is $0.95/1M input on Standard versus $1.50 on Priority). Fast is a higher-speed path targeting 100+ tokens/sec on the same model at roughly double the Standard rate (Kimi K2.6 Fast is $2.00/1M input). Priority is set with the service_tier parameter; Fast is selected by using a different model ID. As of 2026-07-29 a fourth mechanic, US-only Serverless, prices where a request runs rather than how fast or reliably: routing to US-only infrastructure costs a flat 10% premium over the base model's serverless price on most models (Kimi K3 US is $3.30/1M input versus Kimi K3's $3.00/1M) — though Fireworks has already published one exception, as of 2026-08-11: GLM 5.2 Fast US carries no premium and matches the global GLM 5.2 Fast rate exactly.
Does Fireworks have a free tier, and is Fire Pass really free? ▼
There are two different free entry points. New accounts receive $1 in trial credits — enough to confirm the API works, not enough to evaluate a real workload. Fire Pass is the other one, and it genuinely costs nothing per token on the open-weight models it includes; it is activated with a promo code rather than a payment method. The constraints are the price: Fire Pass is limited to personal development and agentic coding harnesses and is explicitly prohibited for production workloads, with violations able to result in the pass being revoked, and it runs on a dedicated fpk_ API key so every other model still bills against your normal Fireworks key. Fireworks also labels it experimental, with features, availability and pricing subject to change — so it is a good way to build a habit, not a foundation for a product.
Is this Fireworks AI pricing calculator free and accurate? ▼
Yes — it's completely free with no signup. It uses the latest publicly available Fireworks AI pricing (verified August 2026). Actual costs may vary with volume discounts or enterprise terms — always confirm on Fireworks AI's official pricing page.