Ask
Packaging

Modal restricts Shared API token-based pricing to Team and Enterprise plans

Modal pricing

Modal now limits its token-billed Shared API (launched with Kimi K3) to Team and Enterprise plans; Starter keeps the same models via a per-second Auto Endpoint instead.

Before

Modal's Shared API launch post said the OpenAI-compatible, token-billed endpoint was covered by Starter's standard $30/month free-compute offer, implying access on any plan.

After

The same post now says the Shared API's token-based pricing is available to Team and Enterprise customers only; Starter and other plans retain access to the same hosted models (e.g. Kimi K3) through a dedicated per-second Auto Endpoint instead.

Modal’s July 29 announcement of an OpenAI-compatible Shared API for Kimi K3 originally framed it as available on any plan, with Starter’s $30/month free-compute credit covering ongoing usage. As of this week, the same blog post has been edited to gate Shared API token-based pricing to Team ($250/mo + compute) and Enterprise customers only. Starter and other lower tiers are not locked out of the underlying model — they can still reach Kimi K3 through Modal’s per-second Auto Endpoint — but lose access to the token-metered pricing surface itself. Modal has not published per-token input/output rates for the Shared API on its pricing page or billing docs.

From Modal's pricing timeline
Shared API Access Narrowed to Team and Enterprise

Modal edited its Kimi K3 / Shared API announcement post to restrict the token-based Shared API to Team and Enterprise customers only, removing earlier language that had implied Starter's $30/month credit covered ongoing Shared API usage on any plan. Starter and other plans keep access to the same hosted models via a dedicated per-second Auto Endpoint instead; no per-second rates or plan fees changed.

About Modal
modal.com ↗

Modal runs a per-second pure-usage compute model: GPU rates from T4 at $0.000164/sec to B300 at $0.001972/sec, with B200 at $0.001736/sec, H200 at $0.001261/sec, H100 at $0.001097/sec, RTX PRO 6000 at $0.000842/sec, A100 80GB at $0.000694/sec, A100 40GB at $0.000583/sec, L40S at $0.000542/sec, A10 at $0.000306/sec, and L4 at $0.000222/sec.

Free tier
Yes
Commits
Available
Transparency
public

Modal pricing history

  1. Aug 2026
    Shared API Access Narrowed to Team and Enterprise
  2. Aug 2026
    Team Plan Container Concurrency Raised 5x to 5,000
  3. Jul 2026
    Shared API Launched — Token-Based Pricing Alongside Per-Second GPU
  4. Jul 2026
    B300 GPU + Sandbox/Notebooks Pricing + Rate Modifiers
  5. Feb 2026
    AWS + GCP Marketplace Billing for Enterprise
Full Modal timeline

More Modal activity

All pricing activity