Ask
Packaging

Fireworks adds two new serverless models, including its cheapest input rate yet

Fireworks AI pricing

Fireworks added Muse Glimmer 30B and NVIDIA Nemotron 3.5 Lightning 30B A3B to its serverless rate card; Nemotron 3.5 Lightning's $0.05 input rate is now the cheapest published on the entire card.

Before

Cheapest published serverless input rate was OpenAI GPT OSS 20B at $0.07 per 1M tokens; the model catalog did not include Muse Glimmer or any Nemotron 3.5 variant.

After

NVIDIA Nemotron 3.5 Lightning 30B A3B is priced at $0.05 / $0.01 / $0.20 per 1M tokens (Standard-only), the new cheapest input rate on the card. Muse Glimmer 30B is priced at $0.35 / $0.04 / $1.50 Standard and $0.525 / $0.06 / $2.25 Priority.

Proof of change

Fireworks AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Aug 13, 2026 Captured Aug 14, 2026 · 1 days apart
serverless-pricing
$0.35 $2.25 $0.05

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Aug 13, 2026
Captured Aug 14, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Fireworks’ serverless pricing docs picked up two new priced model rows in this capture, with no other changes elsewhere on the pricing or docs surfaces.

Muse Glimmer 30B joins the card at $0.35 / $0.04 / $1.50 per 1M tokens (input / cached input / output) on the Standard path, with a Priority path at $0.525 / $0.06 / $2.25 — a flat 1.5x uplift consistent with several other models on the card.

NVIDIA Nemotron 3.5 Lightning 30B A3B is the more notable addition: priced at $0.05 / $0.01 / $0.20 per 1M tokens on Standard only (no Priority path published), its $0.05 input rate undercuts the prior cheapest published rate on the card, OpenAI GPT OSS 20B at $0.07. This continues Fireworks’ pattern of expanding its open-weight model catalog with granular, size-appropriate pricing rather than a flat rate across all hosted models.

From Fireworks AI's pricing timeline
Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning

Fireworks quietly added two new rows to its serverless per-model rate card: Muse Glimmer 30B ($0.35 / $0.04 / $1.50 Standard, $0.525 / $0.06 / $2.25 Priority — a flat 1.5x uplift) and NVIDIA Nemotron 3.5 Lightning 30B A3B ($0.05 / $0.01 / $0.20, Standard-only). Nemotron 3.5 Lightning's $0.05 input rate is now the cheapest published on the entire serverless card, undercutting the previous low of OpenAI GPT OSS 20B at $0.07. No other pricing page or docs surface changed in this capture.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity