Ask
Price change

Novita AI: three Time Limited Free LLMs graduate to paid pricing

Novita AI pricing

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash moved off Novita AI's introductory $0 rate to paid per-token pricing, while a new model, Ling 3.0 Tiny, took over the free-launch slot.

Before

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash listed at $0/M input and output under a Time Limited Free tag on Novita's serverless model catalog.

After

Macaron V1 Venti $1.5/M in ($0.3/M cache read) / $4.5/M out; Macaron V1 Tall $0.45/M in ($0.08/M cache read) / $2.6/M out; Ling 3.0 Flash $0.06/M in ($0.012/M cache read) / $0.18/M out. Ling 3.0 Tiny added as the new $0 Time Limited Free listing.

Proof of change

Novita AI's pricing pages, as we captured them on two dates.

Capture only
Captured Aug 4, 2026 Captured Aug 11, 2026 · 7 days apart
models
$4.5 $0.45 $2.6

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Aug 4, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 2 other pages we compared
main-serverless-endpoints
$4.5 $2.6
Captured Aug 4, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

gpu-instances
Captured Aug 4, 2026
Captured Aug 11, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • These prices come from our capture alone — they were not confirmed against an independent second source.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Novita AI’s serverless model catalog (novita.ai/en/pricing and novita.ai/en/models) confirms that three LLMs previously tagged Time Limited Free — Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash — converted to paid per-token pricing by the 2026-08-11 capture, exactly the expiry the promotional tag implied but never dated. Macaron V1 Venti now bills $1.5/M input ($0.3/M cache read) and $4.5/M output; Macaron V1 Tall bills $0.45/M input ($0.08/M cache read) and $2.6/M output; Ling 3.0 Flash bills $0.06/M input ($0.012/M cache read) and $0.18/M output.

A new small model, Ling 3.0 Tiny, was added at $0/M flat under the same Time Limited Free tag, continuing Novita’s pattern of using a free listing as a launch ramp for a new model rather than a permanent tier — the same mechanic seen when Tencent’s Hy3 graduated off free pricing on 2026-07-21. The overall catalog held steady at 177 models, and this capture cycle found no other rate changes: GPU instances, dedicated endpoints, bare-metal nodes, and Agent Sandbox pricing were all byte-identical to the 2026-08-04 capture.

From Novita AI's pricing timeline
Three free LLMs graduate to paid pricing; Ling 3.0 Tiny takes the $0 slot

Macaron V1 Venti, Macaron V1 Tall, and Ling 3.0 Flash moved off their introductory Time Limited Free tag to paid per-token rates (Macaron V1 Venti $1.5/M in · $4.5/M out; Macaron V1 Tall $0.45/M in · $2.6/M out; Ling 3.0 Flash $0.06/M in · $0.18/M out), while a new small model, Ling 3.0 Tiny, was added as the new $0 free-launch listing under the same tag. Catalog held at 177 models; GPU, dedicated-endpoint, bare-metal, and Agent Sandbox rates were unchanged.

About Novita AI
novita.ai ↗

Novita AI is a pay-as-you-go AI cloud offering inference across 174 listed models as of the 2026-08-14 capture (down from 177 at 2026-08-04/2026-08-11, after that 2026-08-04 refresh added Deepseek V3.2, Deepseek V4 Flash, DeepSeek-OCR 2, and PaddleOCR-VL, following a 2026-07-29 delisting of legacy video/image SKUs including Hunyuan Video Fast, PixVerse V4.5, and the Kling-o1 lineup), on-demand and bare-metal GPUs, and secure per-second agent sandboxes under a single API; a 2026-08-25 capture cut the Image catalog from 14 to 5 SKUs and pruned most legacy Video lines while the site's own model-catalog counter read 144, a divergence from the 174 figure this page has tracked that remains unreconciled.

Free tier
Yes
Commits
None
Transparency
public

Novita AI pricing history

  1. Sep 2026
    RTX 5090 cut 23%; H100 pulled from self-serve again; Image Endpoints and sandbox credit removed
  2. Sep 2026
    Macaron V1 Venti and Tall cut 45% across input, output, and cache-read
  3. Aug 2026
    H100 SXM and L40S return to self-serve GPU instances; RTX 6000 Ada and RTX 5090 HF variant removed
  4. Aug 2026
    Deepseek V4 Flash 0731 repriced sharply; RTX 6000 Ada returns; Image/Video catalog cut hard
  5. Aug 2026
    Ling 3.0 Tiny delisted after 3 days, emptying the free-LLM shelf
Full Novita AI timeline

More Novita AI activity

All pricing activity