Ask
Packaging

Fireworks adds a GB300 GPU tier and a region-restricted deployment premium

Fireworks AI pricing

Fireworks added GB300 288 GB to its on-demand GPU card at $18.00/hr and introduced a 1.5x region-restricted deployment premium for US/Europe-pinned dedicated GPUs, mirroring its existing 10% US-only Serverless token premium.

Before

On-demand dedicated GPU card topped out at B300 288 GB ($12.00/hr); no published geographic-routing premium existed for dedicated deployments.

After

GB300 288 GB is now listed at $18.00/hr, the highest published on-demand rate. Region-restricted deployments (US, Europe) are priced at a flat 1.5x the standard on-demand rate and require a Contact Sales request.

Proof of change

Fireworks AI's pricing pages, as we captured them on two dates.

Capture only
Captured Jul 29, 2026 Captured Aug 11, 2026 · 13 days apart
main
$18.00
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 5 other pages we compared
account-quotas
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

docs
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

fire-pass
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

serverless-pricing
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

serving-paths
Captured Jul 29, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • These prices come from our capture alone — they were not confirmed against an independent second source.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Fireworks’ on-demand dedicated GPU pricing page picked up a new hardware tier and a new pricing axis in the same capture. GB300 288 GB joins H100/H200 ($7.00/hr), B200 ($10.00/hr) and B300 ($12.00/hr) at $18.00/hr — the largest published VRAM-per-GPU option and the highest on-demand rate on the card.

Alongside it, the page now documents region-restricted deployments: GPUs pinned to US-only or Europe-only infrastructure are priced at a flat 1.5x the standard on-demand rate, gated behind a Contact Sales request rather than self-serve checkout. This is the first geographic-routing premium Fireworks has published for dedicated GPU deployments, and it parallels the 10% US-only Serverless premium the company added to its token-based serverless card in July 2026 — extending the same geography-as-a-meter logic from per-token inference to per-GPU-hour deployments.

Separately, the docs clarified that the long-published $50/$500/$5,000/$50,000 monthly spend-tier ceilings apply only to legacy self-serve postpaid accounts; today’s default prepaid accounts set an independent monthly spend limit via firectl quota update monthly-spend-usd, while the tier itself continues to gate serverless TPM ceilings and training-GPU allocation regardless of billing mode.

From Fireworks AI's pricing timeline
GB300 GPU tier + region-restricted deployment premium added

Fireworks added GB300 288 GB to the on-demand dedicated GPU card at $18.00/hr, above B300's $12.00/hr — the highest published on-demand rate on the card. The on-demand pricing page also now documents a region-restricted deployments option (US, Europe) priced at a flat 1.5x the standard on-demand rate, requiring a Contact Sales request — the first geographic-routing premium on the dedicated-GPU side of the product, mirroring the existing 10% US-only Serverless premium on the token side. Separately, the docs now note GLM 5.2 Fast US is exempt from that 10% token-side premium (priced identically to global GLM 5.2 Fast), and clarify that the long-published $50/$500/$5,000/$50,000 spend-tier ceilings apply only to legacy self-serve postpaid accounts, not today's default prepaid accounts.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity