Price change

W&B reprices DeepSeek V4-Flash 14x on Serverless Inference

Weights & Biases pricing

Weights & Biases raised its cheapest Serverless Inference model, DeepSeek V4-Flash, from /bin/bash.01//bin/bash.01 to /bin/bash.14 input / /bin/bash.28 output per 1M tokens, adding a cached tier.

Before

DeepSeek V4-Flash: /bin/bash.01 per 1M input / /bin/bash.01 per 1M output, no cached-token rate.

After

DeepSeek V4-Flash: /bin/bash.14 per 1M input / /bin/bash.07 cached / /bin/bash.28 per 1M output.

Weights & Biases repriced the entry-level model in its Serverless Inference catalog. DeepSeek V4-Flash, an experimental 1M-context MoE model, moved from a flat /bin/bash.01 input / /bin/bash.01 output per 1M tokens to /bin/bash.14 input, /bin/bash.07 cached input, and /bin/bash.28 output — a 14x increase on input and 28x on output, plus a new discounted cached-input tier that the old two-part rate did not have.

The move lifts the floor of the whole catalog: the cheapest published input rate on W&B Inference is now /bin/bash.05 per 1M tokens (OpenAI GPT OSS 20B, IBM Granite 4.1 8B, JetBrains Mellum2 12B), where DeepSeek V4-Flash previously undercut everything at /bin/bash.01. The /bin/bash.01 rate read as a loss-leading introductory price on an “Experimental”-badged model; normalising it to /bin/bash.14//bin/bash.28 puts V4-Flash roughly in line with the rest of the cheap tier while still sitting well below DeepSeek V4-Pro at .74 input / .46 output.

Nothing else moved. The cloud pricing page is unchanged from the prior capture: Free at /bin/bash/mo, Pro from 0/month billed monthly, custom Enterprise, storage overage at /bin/bash.03/GB, Weave data ingestion overage at /bin/bash.10/MB, the /mo Pro Inference credit, CoreWeave Sandboxes credits (0/mo Free, 5/mo Pro), and ARIA still free “for a limited time” on Free and Pro.

From Weights & Biases's pricing timeline
DeepSeek V4-Flash repriced 14x on Serverless Inference

The entry-level DeepSeek V4-Flash model went from $0.01 input / $0.01 output per 1M tokens to $0.14 input / $0.07 cached / $0.28 output, adding a cached-token tier and lifting the Inference catalog's cheapest input rate to $0.03 (OpenAI GPT OSS 20B). Cloud tiers (Free / Pro $60 / Enterprise), storage ($0.03/GB) and Weave ingestion ($0.10/MB) were unchanged.

About Weights & Biases
wandb.ai ↗

Weights & Biases (W&B) is the MLOps experiment-tracking standard plus W&B Weave (LLM observability/evals), Models registry, and Serverless Inference — used by OpenAI, Meta, and Toyota across 1M+ users.

Free tier
Yes
Commits
Available
Transparency
public

Weights & Biases pricing history

  1. Jul 2026
    DeepSeek V4-Flash repriced 14x on Serverless Inference
  2. Jul 2026
    ARIA token-priced product added to the pricing page
  3. Jan 2026
    Restructured tiers: storage + ingestion + per-token Inference
  4. May 2025
    CoreWeave completes ~$1.7B acquisition
  5. Jan 2024
    W&B Weave launches for LLM observability
Full Weights & Biases timeline

More Weights & Biases activity

All pricing activity