Ask
Price change

W&B cuts DeepSeek V4-Pro rates ~34%/27%, adds V4-Flash-0731

Weights & Biases pricing

W&B cut DeepSeek V4-Pro Serverless Inference rates roughly 34% on input and 27% on output, and added a new model, DeepSeek V4-Flash-0731, at $0.13 in / $0.28 out per 1M tokens.

Before

DeepSeek V4-Pro $1.74 in / $0.14 cached / $3.48 out per 1M tokens; 31 models on the rate card.

After

DeepSeek V4-Pro $1.15 in / $0.20 cached / $2.55 out per 1M tokens; DeepSeek V4-Flash-0731 added at $0.13 in / $0.07 cached / $0.28 out; 32 models on the rate card.

Weights & Biases’s Serverless Inference rate card moved for the third time in three weeks. DeepSeek V4-Pro — one of the catalog’s pricier flagship models — was cut roughly 34% on input ($1.74 to $1.15 per 1M tokens) and 27% on output ($3.48 to $2.55), while its cached-input rate rose about 43% ($0.14 to $0.20). A new model, DeepSeek V4-Flash-0731, joined the catalog at $0.13 input / $0.07 cached / $0.28 output per 1M tokens, taking the roster from 31 to 32 models.

The published cloud tiers (Free $0, Pro from $60/mo, custom Enterprise), storage overage ($0.03/GB), Weave data ingestion overage ($0.10/MB), and ARIA’s free-for-now token pricing were all unchanged from the July 29 capture. The two vendor surfaces — the canonical Token-Based Pricing rate card and the marketing “Available models” gallery — agree on price for every model they both list, including this move, but the marketing gallery still omits Z.AI GLM 5 and has newly dropped Microsoft Phi 4 Mini 3.8B (which remains priced, unchanged, on the rate card).

From Weights & Biases's pricing timeline
DeepSeek V4-Pro cut ~34%/27%; DeepSeek V4-Flash-0731 added

DeepSeek V4-Pro was repriced from $1.74 input / $0.14 cached / $3.48 output to $1.15 input / $0.20 cached / $2.55 output per 1M tokens (roughly -34% input, -27% output, +43% cached), and a new model, DeepSeek V4-Flash-0731, joined the Serverless Inference catalog at $0.13 in / $0.07 cached / $0.28 out (32 models total). Cloud tiers (Free / Pro $60 / Enterprise), storage ($0.03/GB), and Weave ingestion ($0.10/MB) were unchanged.

About Weights & Biases
wandb.ai ↗

Weights & Biases (W&B) is the MLOps experiment-tracking standard plus W&B Weave (LLM observability/evals), Models registry, and Serverless Inference — used by OpenAI, Meta, and Toyota across 1M+ users.

Free tier
Yes
Commits
Available
Transparency
public

Weights & Biases pricing history

  1. Aug 2026
    DeepSeek V4-Pro cut ~34%/27%; DeepSeek V4-Flash-0731 added
  2. Jul 2026
    Inference rate card repriced again; MiniMax M3 added
  3. Jul 2026
    DeepSeek V4-Flash repriced 14x on Serverless Inference
  4. Jul 2026
    ARIA token-priced product added to the pricing page
  5. Jan 2026
    Restructured tiers: storage + ingestion + per-token Inference
Full Weights & Biases timeline

More Weights & Biases activity

All pricing activity