Ask
Packaging

DeepInfra adds a Flex service tier at 0.8x the base per-token rate

DeepInfra pricing

DeepInfra added a third per-request service tier: Flex bills 0.8x base price for non-production and asynchronous work, joining Standard (1x) and Priority (1.5x).

Before

Two per-request service tiers: Standard at 1x base price (default) and Priority at 1.5x base price for faster time-to-first-token.

After

Three per-request service tiers: Flex at 0.8x base price (slower responses, occasional unavailability), Standard at 1x, and Priority at 1.5x.

Proof of change

DeepInfra's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 14, 2026 Captured Jul 21, 2026 · 7 days apart
enterprise
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 3 other pages we compared
main
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

models
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

startup-credits
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

DeepInfra has extended the per-request Service Tier control it launched at the end of June into a three-point cost/latency ladder. The new Flex tier bills at 0.8x base price and is described on the pricing page as “lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability.” It sits below the default Standard tier (1x base price, best-effort scheduling) and the Priority tier (1.5x base price, scheduled ahead of standard traffic for faster time-to-first-token).

Flex is not a page-copy-only change: it also appears as a new browsable filter in the model directory alongside Priority, and individual model cards now carry Flex and Priority badges showing which tiers each model supports. Having spent 2024-2025 competing almost entirely on headline rate cuts, DeepInfra now has a two-sided quality-of-service ladder around its base price — a discount lane for batch and dev traffic and a premium lane for latency-sensitive workloads — without adding a plan, a seat, or a contract.

Headline rates are otherwise unchanged this capture: DeepSeek-V3.1 stays at $0.25 in / $0.95 out per 1M; dedicated GPU-hour rates hold at A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69 and B300 $4.89; on-demand instances run 1xB200 $3.69/hr to 8xB200 $29.52/hr; DeepCluster stays at $2.99/GPU-hr (3-year) and $1.98/GPU-hr (5-year); and usage tiers still span $20 to $10,000 invoicing thresholds. The one rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M.

From DeepInfra's pricing timeline
Flex service tier added at 0.8× base price

DeepInfra adds a third per-request Service Tier: Flex, priced at 0.8× base price and described as "lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability" — joining Standard (1×) and Priority (1.5×) and turning the service-tier control into a three-point cost/latency ladder. Flex also appears as a new capability filter in the model directory and as a per-model badge on model cards. Headline rates are otherwise unchanged (DeepSeek-V3.1 $0.25/$0.95, A100 $0.89 / H100 $2.20 / H200 $2.69 / B200 $3.69 / B300 $4.89 per GPU-hour, 8×B200 $29.52/hr, DeepCluster $1.98–$2.99/GPU-hr, usage tiers $20–$10,000); the only rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M (source: deepinfra.com/pricing 2026-07-21).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart. DeepSeek-V4-Flash-0731 was cut again on 2026-09-07 in the /pricing page's DeepSeek table, from $0.08 to $0.06 per 1M input tokens (output flat at $0.18) — its second tracked cut, after the 2026-08-11 move from $0.09.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Sep 2026
    Claude catalog restored to seven models; DeepSeek-V4-Flash-0731 cut again
  2. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  3. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  4. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  5. Jul 2026
    Flex service tier added at 0.8× base price
Full DeepInfra timeline

More DeepInfra activity

All pricing activity