Ask
Price change

DeepInfra raises GPU-hour rates and DeepSeek-V3.1 token price

DeepInfra pricing

DeepInfra lifted dedicated GPU rates 16-32% (H100 to $2.20, B200 to $3.69, B300 to $4.89/GPU-hr) and raised DeepSeek-V3.1 to $0.25/$0.95 per 1M tokens, a notable reversal for a vendor known for cutting prices.

Before

Custom-LLM GPUs: H100 $1.79, H200 $2.19, B200 $2.79, B300 $4.20/GPU-hr; on-demand 1xB200 $2.79/hr. DeepSeek-V3.1 $0.21 in / $0.79 out per 1M.

After

Custom-LLM GPUs: H100 $2.20, H200 $2.69, B200 $3.69, B300 $4.89/GPU-hr (A100 flat at $0.89); on-demand 1xB200 $3.69/hr, 8xB200 $29.52/hr. DeepSeek-V3.1 $0.25 in / $0.95 out per 1M.

DeepInfra, whose reputation in cost-sensitive communities was built on relentless public price cuts, raised its dedicated GPU-hour rates for the first time across the tracked span. Custom-LLM rates moved H100 $1.79 to $2.20 (+23%), H200 $2.19 to $2.69 (+23%), B200 $2.79 to $3.69 (+32%), and B300 $4.20 to $4.89 (+16%); only the A100 held at $0.89/GPU-hour. On-demand B200 instances moved in lockstep — a single card went from $2.79 to $3.69/hr and an 8-card box from $22.32 to $29.52/hr.

On the per-token side, the popular DeepSeek-V3.1 rose to $0.25 in / $0.95 out per 1M (from $0.21 / $0.79), while Llama-3.1-8B-Turbo output edged down to $0.03. The model catalog also expanded (DeepSeek-V4-Flash $0.09/$0.18, gemini-3.x, gemma-4, claude-fable-5 $10/$50, claude-sonnet-5 $2/$10). DeepCluster reserved pricing ($1.98-$2.99/GPU-hr), the $20-$10,000 usage tiers, and the 1.5x Priority service tier were unchanged.

From DeepInfra's pricing timeline
GPU-hour rates raised; DeepSeek-V3.1 token price up

DeepInfra raises its dedicated GPU-hour rates for the first time in the tracked span: custom-LLM H100 $1.79→$2.20, H200 $2.19→$2.69, B200 $2.79→$3.69, B300 $4.20→$4.89 per GPU-hour (A100 holds at $0.89); on-demand B200 instances move in lockstep (1×B200 $2.79→$3.69/hr, 8×B200 $22.32→$29.52/hr). DeepSeek-V3.1 per-token rises to $0.25 in / $0.95 out (from $0.21 / $0.79) and Llama-3.1-8B-Turbo output drops to $0.03. The model catalog expands (DeepSeek-V4-Flash $0.09/$0.18, gemini-3.x, gemma-4, claude-fable-5 $10/$50, claude-sonnet-5). DeepCluster ($1.98–$2.99/GPU-hr), usage tiers ($20–$10,000), and the Priority service tier (1.5×) are unchanged (source: deepinfra.com/pricing & /gpu-instances 2026-07-14).

About DeepInfra
deepinfra.com ↗

DeepInfra is a serverless inference cloud that bills per-token for language and embedding models and per-inference-execution-time for most other models, with no contracts or upfront costs. Representative per-1M-token rates: DeepSeek-V3.1 $0.25 in / $0.95 out, DeepSeek-V4-Pro $1.30 / $2.60, Llama-3.3-70B-Turbo $0.10 / $0.32, Llama-3.1-8B $0.02 / $0.04. Llama-3.1-8B-Instruct-Turbo's output rate rose 33% on 2026-07-29 (from $0.03), DeepInfra's second token-price increase since its 2026-07-14 reversal, while gemma-4-31B-it-turbo was cut to $0.09 in / $0.34 out per 1M in the same update. GLM-5.2, Z-AI's flagship long-horizon model featured on DeepInfra's /models catalog, was then cut about 20% on 2026-08-04 (from $0.93 in / $3.00 out to $0.75 / $2.40 per 1M), the first outright cut since the July reversal began, though its price never appears on the main /pricing page. GLM-5.2 was cut again on 2026-08-28, this time via a new 35%-off promotional tag taking it from $0.75 / $2.40 to $0.488 / $1.56 per 1M — still only visible on /models and /deepstart.

Free tier
No
Commits
Available
Transparency
public

DeepInfra pricing history

  1. Aug 2026
    Six of seven Claude models pulled from the /pricing page; GLM-5.2 gets a new 35% promo
  2. Aug 2026
    GLM-5.2 cut ~20% — but only on the /models catalog
  3. Jul 2026
    Second rate rise: Llama-3.1-8B-Instruct-Turbo output up 33%
  4. Jul 2026
    Flex service tier added at 0.8× base price
  5. Jul 2026
    GPU-hour rates raised; DeepSeek-V3.1 token price up
Full DeepInfra timeline

More DeepInfra activity

All pricing activity