Modal launches Shared API — token-based pricing alongside per-second GPU billing
Modal added a new OpenAI-compatible Shared API billed by token rather than GPU-second, launching first with Moonshot's Kimi K3 model; per-token rates are not yet published on Modal's pricing page.
Modal billed only by GPU/CPU/memory-second (plus flat Starter/Team/Enterprise plan fees) across all compute products.
A parallel Shared API product now bills by token for select hosted models (starting with Kimi K3), alongside the existing per-second GPU rate card; Starter's $30/month free credit applies to Shared API usage too.
Modal’s pricing page began surfacing a banner (“Kimi K3 is live. Try the new Shared API with token-based pricing”) pointing to a new product line: an OpenAI-compatible Shared API endpoint metered by token rather than by GPU-second. This is Modal’s first departure from pure per-second compute billing since founding — it runs alongside, not instead of, the existing GPU/CPU/memory rate card, giving customers a choice between dedicated per-second GPU capacity and shared per-token inference for supported models.
The first model available on the Shared API is Moonshot’s Kimi K3. Modal has not yet published a per-token rate card on its public pricing page or billing docs, so exact input/output token prices are unknown as of this capture. A follow-up discovery pass is needed once Modal publishes the rate card.
Modal launched an OpenAI-compatible Shared API metered by token rather than by GPU-second, going live first with Moonshot's Kimi K3 model. The Shared API runs in parallel with — not instead of — the existing per-second GPU/CPU/memory rate card, and Starter's $30/month free credit grant applies to Shared API usage too. This is Modal's first departure from pure per-second compute billing since founding; per-token input/output rates were not yet published on modal.com/pricing as of this capture.