Modal launches Shared API — token-based pricing alongside per-second GPU billing
Modal added a new OpenAI-compatible Shared API billed by token rather than GPU-second, launching first with Moonshot's Kimi K3 model; per-token rates are not yet published on Modal's pricing page.
Modal billed only by GPU/CPU/memory-second (plus flat Starter/Team/Enterprise plan fees) across all compute products.
A parallel Shared API product now bills by token for select hosted models (starting with Kimi K3), alongside the existing per-second GPU rate card; Starter's $30/month free credit applies to Shared API usage too.
Proof of change
Modal's own wording, on the page where we captured it.
Kimi K3 is live. Try the new Shared API with token-based pricing
Quoted from the capture below — these words appear on the page as shown, not paraphrased.
Showing the whole page as captured — scroll the panel, or open it at full size.
This is a single observation, not a before/after comparison — the change it describes has no visible transition to photograph, so we show the page that states it instead. The pixels are our own capture, unmodified.
Modal’s pricing page began surfacing a banner (“Kimi K3 is live. Try the new Shared API with token-based pricing”) pointing to a new product line: an OpenAI-compatible Shared API endpoint metered by token rather than by GPU-second. This is Modal’s first departure from pure per-second compute billing since founding — it runs alongside, not instead of, the existing GPU/CPU/memory rate card, giving customers a choice between dedicated per-second GPU capacity and shared per-token inference for supported models.
The first model available on the Shared API is Moonshot’s Kimi K3. Modal has not yet published a per-token rate card on its public pricing page or billing docs, so exact input/output token prices are unknown as of this capture. A follow-up discovery pass is needed once Modal publishes the rate card.
Modal launched an OpenAI-compatible Shared API metered by token rather than by GPU-second, going live first with Moonshot's Kimi K3 model. The Shared API runs in parallel with — not instead of — the existing per-second GPU/CPU/memory rate card, and Starter's $30/month free credit grant applies to Shared API usage too. This is Modal's first departure from pure per-second compute billing since founding; per-token input/output rates were not yet published on modal.com/pricing as of this capture.