Fireworks launches Kimi K3 flagship model and a pay-per-token Serverless Training API
Fireworks added Kimi K3 (plus Fast and a 10%-premium US-only variant) to its serverless rate card, shipped a new Serverless Training API for per-token LoRA training, and switched Fire Pass's free model to Kimi K3 Fast with a 1M-token context window.
Headline serverless models topped out at Kimi K2.7/K2.6 and GLM 5.x; no serverless training product existed; Fire Pass granted free access to GLM 5.2 Fast (256k context).
Kimi K3 leads the rate card at $3.00 / $0.30 / $15.00 per 1M input/cached/output tokens (Standard), with Fast ($4.50/$0.45/$22.50) and a US-only variant priced at a flat 10% premium ($3.30/$0.33/$16.50). A new Serverless Training API prices LoRA training on Qwen 3.5 9B, Qwen 3.6 27B and Kimi K3 by prefill/cached-prefill/sample/train tokens ($0.66-$32.55 per 1M depending on stage and model). Fire Pass now grants Kimi K3 Fast (1M context) instead of GLM 5.2 Fast.
Fireworks used its site banner to promote Kimi K3 as a new flagship model on 2026-07-29, adding it to the serverless price card alongside a Fast variant and a US-only-routing variant. The US variant is the first model to carry its own priced row for a broader mechanic newly documented in the docs: US-only Serverless endpoints are billed at a flat 10% premium over the base model’s serverless price on any model, not just Kimi K3.
The bigger structural change is a new product: the Serverless Training API, a Tinker-compatible offering that attaches to a shared, always-on trainer pool for LoRA training with no provisioning step and no idle cost. It bills per token across four dimensions — prefill, cached prefill, sample, and train — currently covering three models (Qwen 3.5 9B, Qwen 3.6 27B, Kimi K3), with more “coming soon” per the docs.
Fire Pass, the promo-code pass launched a week earlier with zero per-token pricing on one open-weight model, swapped its included model from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context) and added FireConnect auto-detection plus new supported harnesses (Codex, Pi, LangChain Deep Agents). The rest of the rate card is unchanged: H100/H200 dedicated at $7.00/hr, B200 at $10.00/hr, B300 at $12.00/hr, fine-tuning from $0.50 per 1M training tokens, embeddings from $0.008 per 1M, and batch inference at 50% of serverless.
Fireworks added Kimi K3 (plus a Fast variant and a US-only variant) to its serverless rate card, publishing the first priced row for a new US-only Serverless mechanic — a flat 10% premium over base serverless pricing that the docs describe as applying to any model. Fireworks also shipped the Serverless Training API, a Tinker-compatible product that meters LoRA training per token (prefill, cached prefill, sample, train) on a shared trainer pool with no provisioning or idle cost, covering Qwen 3.5 9B, Qwen 3.6 27B, and Kimi K3. Fire Pass's included free model swapped from GLM 5.2 Fast (256k context) to Kimi K3 Fast (1M context).