Ask
Packaging

Fireworks adds two new serverless models, including its cheapest input rate yet

Fireworks AI pricing

Fireworks added Muse Glimmer 30B and NVIDIA Nemotron 3.5 Lightning 30B A3B to its serverless rate card; Nemotron 3.5 Lightning's $0.05 input rate is now the cheapest published on the entire card.

Before

Cheapest published serverless input rate was OpenAI GPT OSS 20B at $0.07 per 1M tokens; the model catalog did not include Muse Glimmer or any Nemotron 3.5 variant.

After

NVIDIA Nemotron 3.5 Lightning 30B A3B is priced at $0.05 / $0.01 / $0.20 per 1M tokens (Standard-only), the new cheapest input rate on the card. Muse Glimmer 30B is priced at $0.35 / $0.04 / $1.50 Standard and $0.525 / $0.06 / $2.25 Priority.

Fireworks’ serverless pricing docs picked up two new priced model rows in this capture, with no other changes elsewhere on the pricing or docs surfaces.

Muse Glimmer 30B joins the card at $0.35 / $0.04 / $1.50 per 1M tokens (input / cached input / output) on the Standard path, with a Priority path at $0.525 / $0.06 / $2.25 — a flat 1.5x uplift consistent with several other models on the card.

NVIDIA Nemotron 3.5 Lightning 30B A3B is the more notable addition: priced at $0.05 / $0.01 / $0.20 per 1M tokens on Standard only (no Priority path published), its $0.05 input rate undercuts the prior cheapest published rate on the card, OpenAI GPT OSS 20B at $0.07. This continues Fireworks’ pattern of expanding its open-weight model catalog with granular, size-appropriate pricing rather than a flat rate across all hosted models.

From Fireworks AI's pricing timeline
Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning

Fireworks quietly added two new rows to its serverless per-model rate card: Muse Glimmer 30B ($0.35 / $0.04 / $1.50 Standard, $0.525 / $0.06 / $2.25 Priority — a flat 1.5x uplift) and NVIDIA Nemotron 3.5 Lightning 30B A3B ($0.05 / $0.01 / $0.20, Standard-only). Nemotron 3.5 Lightning's $0.05 input rate is now the cheapest published on the entire serverless card, undercutting the previous low of OpenAI GPT OSS 20B at $0.07. No other pricing page or docs surface changed in this capture.

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments at $7/hr H100/H200, $10/hr B200, $12/hr B300 — rates that Fireworks has announced will rise to $8/$13/$15 respectively (GB300 $18 to $20) effective September 1, 2026.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  2. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  3. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
  4. Aug 2026
    GB300 GPU tier + region-restricted deployment premium added
  5. Jul 2026
    Kimi K3 flagship, US-only Serverless premium, and Serverless Training API
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity