W&B reprices Serverless Inference catalog, adds MiniMax M3
W&B cut Serverless Inference rates up to 45% on GLM 5.2, Kimi K2.7/K2.6, and Gemma 4 31B; raised GPT OSS 120B output 21%; and added MiniMax M3 to the catalog.
Z.AI GLM 5.2 $1.39 in/$4.40 out; Kimi K2.7 Code $0.94/$4.00; Kimi K2.6 $0.95/$4.00; Gemma 4 31B $0.12/$0.35 (+$0.09 cached); GPT OSS 120B $0.04/$0.14; 30 models total.
Z.AI GLM 5.2 $0.76/$2.42; Kimi K2.7 Code $0.71/$3.50; Kimi K2.6 $0.65/$3.41; Gemma 4 31B $0.10/$0.34 (no cached tier); GPT OSS 120B $0.03/$0.17; MiniMax M3 added; 31 models total.
Weights & Biases’s Serverless Inference rate card — the per-token pricing for its hosted open-weight models — had its second reprice in two weeks. Five models got cheaper by 15-45% per 1M tokens, led by Z.AI GLM 5.2’s roughly 45% cut on both input ($1.39 to $0.76) and output ($4.40 to $2.42); Moonshot AI’s Kimi K2.7 Code and Kimi K2.6 fell 25-32% on input. OpenAI GPT OSS 120B’s input rate dropped to $0.03 while its output rate rose 21% to $0.17, and a new model, MiniMax M3, joined the catalog at $0.23 input / $0.96 output per 1M tokens (31 models total, up from 30).
Google Gemma 4 31B’s discounted cached-input tier ($0.09/1M) also disappeared from both the marketing model list and the canonical Token-Based Pricing rate card. The published cloud tiers (Free $0, Pro from $60/mo, custom Enterprise), storage overage ($0.03/GB), Weave data ingestion overage ($0.10/MB), and ARIA’s free-for-now token pricing were all unchanged from the July 21 capture.
Five more Serverless Inference models were cut 15-45% per 1M tokens (Z.AI GLM 5.2 ~45% to $0.76 in/$2.42 out; Moonshot Kimi K2.7 Code and K2.6 down 25-32%; Google Gemma 4 31B down, losing its cached-input tier), OpenAI GPT OSS 120B's output rose 21% to $0.17, and a new model, MiniMax M3, joined the catalog (31 models total). Cloud tiers, storage, and Weave ingestion were unchanged.