Ask
Price change

Novita AI: three Time Limited Free LLMs graduate to paid pricing

Novita AI pricing

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash moved off Novita AI's introductory $0 rate to paid per-token pricing, while a new model, Ling 3.0 Tiny, took over the free-launch slot.

Before

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash listed at $0/M input and output under a Time Limited Free tag on Novita's serverless model catalog.

After

Macaron V1 Venti $1.5/M in ($0.3/M cache read) / $4.5/M out; Macaron V1 Tall $0.45/M in ($0.08/M cache read) / $2.6/M out; Ling 3.0 Flash $0.06/M in ($0.012/M cache read) / $0.18/M out. Ling 3.0 Tiny added as the new $0 Time Limited Free listing.

Novita AI’s serverless model catalog (novita.ai/en/pricing and novita.ai/en/models) confirms that three LLMs previously tagged Time Limited Free — Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash — converted to paid per-token pricing by the 2026-08-11 capture, exactly the expiry the promotional tag implied but never dated. Macaron V1 Venti now bills $1.5/M input ($0.3/M cache read) and $4.5/M output; Macaron V1 Tall bills $0.45/M input ($0.08/M cache read) and $2.6/M output; Ling 3.0 Flash bills $0.06/M input ($0.012/M cache read) and $0.18/M output.

A new small model, Ling 3.0 Tiny, was added at $0/M flat under the same Time Limited Free tag, continuing Novita’s pattern of using a free listing as a launch ramp for a new model rather than a permanent tier — the same mechanic seen when Tencent’s Hy3 graduated off free pricing on 2026-07-21. The overall catalog held steady at 177 models, and this capture cycle found no other rate changes: GPU instances, dedicated endpoints, bare-metal nodes, and Agent Sandbox pricing were all byte-identical to the 2026-08-04 capture.

From Novita AI's pricing timeline
Three free LLMs graduate to paid pricing; Ling 3.0 Tiny takes the $0 slot

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash moved off their introductory Time Limited Free tag to paid per-token rates (Macaron V1 Venti $1.5/M in · $4.5/M out; Macaron V1 Tall $0.45/M in · $2.6/M out; Ling 3.0 Flash $0.06/M in · $0.18/M out), while a new small model, Ling 3.0 Tiny, was added as the new $0 free-launch listing under the same tag. Catalog held at 177 models; GPU, dedicated-endpoint, bare-metal, and Agent Sandbox rates were unchanged.

About Novita AI
novita.ai ↗

Novita AI is a pay-as-you-go AI cloud offering inference across 174 listed models as of the 2026-08-14 capture (down from 177 at 2026-08-04/2026-08-11, after that 2026-08-04 refresh added Deepseek V3.2, Deepseek V4 Flash, DeepSeek-OCR 2, and PaddleOCR-VL, following a 2026-07-29 delisting of legacy video/image SKUs including Hunyuan Video Fast, PixVerse V4.5, and the Kling-o1 lineup), on-demand and bare-metal GPUs, and secure per-second agent sandboxes under a single API; a 2026-08-25 capture cut the Image catalog from 14 to 5 SKUs and pruned most legacy Video lines while the site's own model-catalog counter read 144, a divergence from the 174 figure this page has tracked that remains unreconciled.

Free tier
Yes
Commits
None
Transparency
public

Novita AI pricing history

  1. Aug 2026
    H100 SXM and L40S return to self-serve GPU instances; RTX 6000 Ada and RTX 5090 HF variant removed
  2. Aug 2026
    Deepseek V4 Flash 0731 repriced sharply; RTX 6000 Ada returns; Image/Video catalog cut hard
  3. Aug 2026
    Ling 3.0 Tiny delisted after 3 days, emptying the free-LLM shelf
  4. Aug 2026
    Three free LLMs graduate to paid pricing; Ling 3.0 Tiny takes the $0 slot
  5. Aug 2026
    RTX 6000 Ada dropped from self-serve GPUs; base RTX 5090 tier added at $0.73/hr
Full Novita AI timeline

More Novita AI activity

All pricing activity