Packaging

DeepInfra adds a Flex service tier at 0.8x the base per-token rate

DeepInfra pricing

DeepInfra added a third per-request service tier: Flex bills 0.8x base price for non-production and asynchronous work, joining Standard (1x) and Priority (1.5x).

Before

Two per-request service tiers: Standard at 1x base price (default) and Priority at 1.5x base price for faster time-to-first-token.

After

Three per-request service tiers: Flex at 0.8x base price (slower responses, occasional unavailability), Standard at 1x, and Priority at 1.5x.

DeepInfra has extended the per-request Service Tier control it launched at the end of June into a three-point cost/latency ladder. The new Flex tier bills at 0.8x base price and is described on the pricing page as “lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability.” It sits below the default Standard tier (1x base price, best-effort scheduling) and the Priority tier (1.5x base price, scheduled ahead of standard traffic for faster time-to-first-token).

Flex is not a page-copy-only change: it also appears as a new browsable filter in the model directory alongside Priority, and individual model cards now carry Flex and Priority badges showing which tiers each model supports. Having spent 2024-2025 competing almost entirely on headline rate cuts, DeepInfra now has a two-sided quality-of-service ladder around its base price — a discount lane for batch and dev traffic and a premium lane for latency-sensitive workloads — without adding a plan, a seat, or a contract.

Headline rates are otherwise unchanged this capture: DeepSeek-V3.1 stays at $0.25 in / $0.95 out per 1M; dedicated GPU-hour rates hold at A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69 and B300 $4.89; on-demand instances run 1xB200 $3.69/hr to 8xB200 $29.52/hr; DeepCluster stays at $2.99/GPU-hr (3-year) and $1.98/GPU-hr (5-year); and usage tiers still span $20 to $10,000 invoicing thresholds. The one rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M.

From DeepInfra's pricing timeline
Flex service tier added at 0.8× base price

DeepInfra adds a third per-request Service Tier: Flex, priced at 0.8× base price and described as "lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability" — joining Standard (1×) and Priority (1.5×) and turning the service-tier control into a three-point cost/latency ladder. Flex also appears as a new capability filter in the model directory and as a per-model badge on model cards. Headline rates are otherwise unchanged (DeepSeek-V3.1 $0.25/$0.95, A100 $0.89 / H100 $2.20 / H200 $2.69 / B200 $3.69 / B300 $4.89 per GPU-hour, 8×B200 $29.52/hr, DeepCluster $1.98–$2.99/GPU-hr, usage tiers $20–$10,000); the only rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M (source: deepinfra.com/pricing 2026-07-21).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.03.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Jul 2026
    Flex service tier added at 0.8× base price
  2. Jul 2026
    GPU-hour rates raised; DeepSeek-V3.1 token price up
  3. Jun 2026
    Priority Service Tier added (1.5× per-token multiplier)
  4. Jun 2026
    Per-token + per-GPU-hour + DeepCluster reserved capacity
  5. May 2026
    DeepCluster (customer-owned B300) and $107M Series B
Full DeepInfra timeline

More DeepInfra activity

All pricing activity