Ask
Launch

Fireworks launches Kimi K3 flagship model and a pay-per-token Serverless Training API

Fireworks AI pricing

Fireworks added Kimi K3 (plus Fast and a 10%-premium US-only variant) to its serverless rate card, shipped a new Serverless Training API for per-token LoRA training, and switched Fire Pass's free model to Kimi K3 Fast with a 1M-token context window.

Before

Headline serverless models topped out at Kimi K2.7/K2.6 and GLM 5.x; no serverless training product existed; Fire Pass granted free access to GLM 5.2 Fast (256k context).

After

Kimi K3 leads the rate card at $3.00 / $0.30 / $15.00 per 1M input/cached/output tokens (Standard), with Fast ($4.50/$0.45/$22.50) and a US-only variant priced at a flat 10% premium ($3.30/$0.33/$16.50). A new Serverless Training API prices LoRA training on Qwen 3.5 9B, Qwen 3.6 27B and Kimi K3 by prefill/cached-prefill/sample/train tokens ($0.66-$32.55 per 1M depending on stage and model). Fire Pass now grants Kimi K3 Fast (1M context) instead of GLM 5.2 Fast.

Proof of change

Fireworks AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 22, 2026 Captured Jul 29, 2026 · 7 days apart
main
$0.66 $0.13 $1.99 $1.46 $1.86 $0.37

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 5 other pages we compared
serverless-pricing
$3.00 $15.00 $3.75 $0.37 $18.75 $4.50

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

account-quotas
Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

docs
Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

fire-pass
Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

serving-paths
Captured Jul 22, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Fireworks used its site banner to promote Kimi K3 as a new flagship model on 2026-07-29, adding it to the serverless price card alongside a Fast variant and a US-only-routing variant. The US variant is the first model to carry its own priced row for a broader mechanic newly documented in the docs: US-only Serverless endpoints are billed at a flat 10% premium over the base model’s serverless price on any model, not just Kimi K3.

The bigger structural change is a new product: the Serverless Training API, a Tinker-compatible offering that attaches to a shared, always-on trainer pool for LoRA training with no provisioning step and no idle cost. It bills per token across four dimensions — prefill, cached prefill, sample, and train — currently covering three models (Qwen 3.5 9B, Qwen 3.6 27B, Kimi K3), with more “coming soon” per the docs.

Fire Pass, the promo-code pass launched a week earlier with zero per-token pricing on one open-weight model, swapped its included model from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context) and added FireConnect auto-detection plus new supported harnesses (Codex, Pi, LangChain Deep Agents). The rest of the rate card is unchanged: H100/H200 dedicated at $7.00/hr, B200 at $10.00/hr, B300 at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.

From Fireworks AI's pricing timeline
Kimi K3 flagship, US-only Serverless premium, and Serverless Training API

Fireworks added Kimi K3 (plus a Fast variant and a US-only variant) to its serverless rate card, publishing the first priced row for a new US-only Serverless mechanic — a flat 10% premium over base serverless pricing that the docs describe as applying to any model. Fireworks also shipped the Serverless Training API, a Tinker-compatible product that meters LoRA training per token (prefill, cached prefill, sample, train) on a shared trainer pool with no provisioning or idle cost, covering Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3. Fire Pass's included free model swapped from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context).

About Fireworks AI
fireworks.ai ↗

Fireworks AI runs a pure-usage inference platform: serverless per-token endpoints led by Kimi K3 (its 2026-07-29 flagship model, from $3.00/1M input on Standard) alongside other open-weight models (DeepSeek, GLM, Qwen), plus on-demand dedicated GPU deployments now at $8/hr H100/H200, $13/hr B200, $15/hr B300, and $20/hr GB300 — a price increase that took effect September 1, 2026 and is confirmed live as of this capture, up from $7/$10/$12/$18 respectively.

Free tier
Yes
Commits
Available
Transparency
public

Fireworks AI pricing history

  1. Sep 2026
    On-demand GPU hike confirmed live; US-only Serverless premium raised to 50%
  2. Sep 2026
    Reserved Throughput launched — sales-led SLA capacity commit
  3. Aug 2026
    DeepSeek V4 Flash (0731) repriced; Qwen 3.8 Max added
  4. Aug 2026
    Two new serverless models added — Muse Glimmer 30B and Nemotron 3.5 Lightning
  5. Aug 2026
    On-demand GPU price increase announced, effective September 1, 2026
Full Fireworks AI timeline

More Fireworks AI activity

All pricing activity