Ask
Price change

W&B reprices DeepSeek V4-Flash 14x on Serverless Inference

Weights & Biases pricing

Weights & Biases raised its cheapest Serverless Inference model, DeepSeek V4-Flash, from $0.01/$0.01 to $0.14 input / $0.28 output per 1M tokens, adding a cached tier.

Before

DeepSeek V4-Flash: $0.01 per 1M input / $0.01 per 1M output, no cached-token rate.

After

DeepSeek V4-Flash: $0.14 per 1M input / $0.07 cached / $0.28 per 1M output.

Proof of change

Weights & Biases's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 14, 2026 Captured Jul 21, 2026 · 7 days apart
inference
$0.01 $0.07 $0.28
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • 2 pages had no counterpart in the earlier capture and are not shown.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Weights & Biases repriced the entry-level model in its Serverless Inference catalog. DeepSeek V4-Flash, an experimental 1M-context MoE model, moved from a flat $0.01 input / $0.01 output per 1M tokens to $0.14 input, $0.07 cached input, and $0.28 output — a 14x increase on input and 28x on output, plus a new discounted cached-input tier that the old two-part rate did not have.

The move lifts the floor of the whole catalog: the cheapest published input rate on W&B Inference is now $0.05 per 1M tokens (OpenAI GPT OSS 20B, IBM Granite 4.1 8B, JetBrains Mellum2 12B), where DeepSeek V4-Flash previously undercut everything at $0.01. The $0.01 rate read as a loss-leading introductory price on an “Experimental”-badged model; normalising it to $0.14/$0.28 puts V4-Flash roughly in line with the rest of the cheap tier while still sitting well below DeepSeek V4-Pro at $1.74 input / $3.46 output.

Nothing else moved. The cloud pricing page is unchanged from the prior capture: Free at $0/mo, Pro from $60/month billed monthly, custom Enterprise, storage overage at $0.03/GB, Weave data ingestion overage at $0.10/MB, the $5/mo Pro Inference credit, CoreWeave Sandboxes credits ($10/mo Free, $25/mo Pro), and ARIA still free “for a limited time” on Free and Pro.

From Weights & Biases's pricing timeline
DeepSeek V4-Flash repriced 14x on Serverless Inference

The entry-level DeepSeek V4-Flash model went from $0.01 input / $0.01 output per 1M tokens to $0.14 input / $0.07 cached / $0.28 output, adding a cached-token tier and lifting the Inference catalog's cheapest input rate to $0.03 (OpenAI GPT OSS 20B). Cloud tiers (Free / Pro $60 / Enterprise), storage ($0.03/GB) and Weave ingestion ($0.10/MB) were unchanged.

About Weights & Biases
wandb.ai ↗

Weights & Biases (W&B) is the MLOps experiment-tracking standard plus W&B Weave (LLM observability/evals), Models registry, and Serverless Inference — used by OpenAI, Meta, and Toyota across 1M+ users.

Free tier
Yes
Commits
Available
Transparency
public

Weights & Biases pricing history

  1. Aug 2026
    DeepSeek V4-Pro cut ~34%/27%; DeepSeek V4-Flash-0731 added
  2. Jul 2026
    Inference rate card repriced again; MiniMax M3 added
  3. Jul 2026
    DeepSeek V4-Flash repriced 14x on Serverless Inference
  4. Jul 2026
    ARIA token-priced product added to the pricing page
  5. Jan 2026
    Restructured tiers: storage + ingestion + per-token Inference
Full Weights & Biases timeline

More Weights & Biases activity

All pricing activity