DeepInfra adds a Flex service tier at 0.8x the base per-token rate
DeepInfra added a third per-request service tier: Flex bills 0.8x base price for non-production and asynchronous work, joining Standard (1x) and Priority (1.5x).
Two per-request service tiers: Standard at 1x base price (default) and Priority at 1.5x base price for faster time-to-first-token.
Three per-request service tiers: Flex at 0.8x base price (slower responses, occasional unavailability), Standard at 1x, and Priority at 1.5x.
Proof of change
DeepInfra's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 3 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
DeepInfra has extended the per-request Service Tier control it launched at the end of June into a three-point cost/latency ladder. The new Flex tier bills at 0.8x base price and is described on the pricing page as “lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability.” It sits below the default Standard tier (1x base price, best-effort scheduling) and the Priority tier (1.5x base price, scheduled ahead of standard traffic for faster time-to-first-token).
Flex is not a page-copy-only change: it also appears as a new browsable filter in the model directory alongside Priority, and individual model cards now carry Flex and Priority badges showing which tiers each model supports. Having spent 2024-2025 competing almost entirely on headline rate cuts, DeepInfra now has a two-sided quality-of-service ladder around its base price — a discount lane for batch and dev traffic and a premium lane for latency-sensitive workloads — without adding a plan, a seat, or a contract.
Headline rates are otherwise unchanged this capture: DeepSeek-V3.1 stays at $0.25 in / $0.95 out per 1M; dedicated GPU-hour rates hold at A100 $0.89, H100 $2.20, H200 $2.69, B200 $3.69 and B300 $4.89; on-demand instances run 1xB200 $3.69/hr to 8xB200 $29.52/hr; DeepCluster stays at $2.99/GPU-hr (3-year) and $1.98/GPU-hr (5-year); and usage tiers still span $20 to $10,000 invoicing thresholds. The one rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M.
DeepInfra adds a third per-request Service Tier: Flex, priced at 0.8× base price and described as "lower cost for non-production and asynchronous work, in exchange for slower responses and occasional unavailability" — joining Standard (1×) and Priority (1.5×) and turning the service-tier control into a three-point cost/latency ladder. Flex also appears as a new capability filter in the model directory and as a per-model badge on model cards. Headline rates are otherwise unchanged (DeepSeek-V3.1 $0.25/$0.95, A100 $0.89 / H100 $2.20 / H200 $2.69 / B200 $3.69 / B300 $4.89 per GPU-hour, 8×B200 $29.52/hr, DeepCluster $1.98–$2.99/GPU-hr, usage tiers $20–$10,000); the only rate move on the model list is Mistral-Nemo-Instruct-2407, cut from $0.02 in / $0.04 out to $0.019 in / $0.03 out per 1M (source: deepinfra.com/pricing 2026-07-21).