Ask
Packaging

DeepInfra adds a Priority Service Tier at 1.5× the base per-token rate

DeepInfra pricing

DeepInfra introduced a per-request Service Tier on its pricing page: Standard stays 1× base price, while a new Priority tier bills 1.5× for faster time-to-first-token during peak demand.

Before

Single per-token rate per model, no scheduling tier — every request priced at the base rate.

After

Per-request Service Tier: Standard (1× base price, default) or Priority (1.5× base price, scheduled ahead of standard traffic for faster TTFT), set via service_tier: "priority".

DeepInfra’s pricing page now exposes a Service Tier control that lets callers trade latency against cost on a per-request basis. The default Standard tier keeps best-effort scheduling at 1× the model’s base per-token price. The new Priority tier schedules a request ahead of standard traffic for faster time-to-first-token during peak demand, billed at 1.5× base price, and is enabled per request by setting service_tier: "priority" (availability varies by model).

This is DeepInfra’s first explicit latency-vs-cost packaging knob layered on top of its pure-usage per-token rates. Headline per-token, per-GPU-hour ($0.89–$4.20/GPU-hr), and DeepCluster ($1.98–$2.99/GPU-hr) rates are otherwise unchanged this capture; the model catalog also expanded (DeepSeek-V4-Flash, Gemini 3.x, Gemma 4, Nemotron-3, Claude Haiku/Sonnet/Opus 4.x) and Qwen3-235B-A22B-Instruct-2507’s input rate edged from $0.071 to $0.09 per 1M tokens.

From DeepInfra's pricing timeline
Priority Service Tier added (1.5× per-token multiplier)

DeepInfra adds a per-request Service Tier control to the pricing page: the default Standard tier bills at 1× base price, while a new Priority tier schedules requests ahead of standard traffic for faster time-to-first-token at 1.5× base price (set via service_tier: "priority"). Headline per-token, per-GPU-hour, and DeepCluster rates are unchanged; the model catalog expands (DeepSeek-V4-Flash, Gemini 3.x, Gemma 4, Nemotron-3, Claude Haiku/Sonnet/Opus 4.x), and Qwen3-235B-A22B-Instruct-2507 input edges to $0.09 (source: deepinfra.com/pricing 2026-06-30).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  2. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  3. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  4. Jul 2026
    Flex service tier added at 0.8× base price
  5. Jul 2026
    GPU-hour rates raised; DeepSeek-V3.1 token price up
Full DeepInfra timeline

More DeepInfra activity

All pricing activity