Baseten adds Kimi K3 and GLM-5.2 Fast, cuts GLM-5.2 cache pricing 46%
Baseten grew its Model APIs catalog to ten models with new flagship Kimi K3 and GLM-5.2 Fast, and cut GLM-5.2's cache-input rate 46% to $0.14 per 1M tokens.
8 Model APIs SKUs: Inkling, GLM-5.2 ($1.40 in / $0.26 cache / $4.40 out), GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.
10 Model APIs SKUs adding Kimi K3 ($3.00 in / $0.30 cache / $15.00 out) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60); GLM-5.2 cache input cut to $0.14 (input/output unchanged at $1.40/$4.40).
Proof of change
Baseten's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 2 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Eight days after pruning its Model APIs catalog from eleven models to eight, Baseten expanded it again — this time by addition rather than retirement. Kimi K3 launched as the platform’s new flagship model at $3.00 input / $0.30 cache input / $15.00 output per 1M tokens, flagged by a homepage “Kimi K3 is here.” banner, and became the most expensive SKU on the rate card by a wide margin (72% above the prior ceiling, DeepSeek V4, on input; over 3x on output). GLM-5.2 Fast joined alongside it at $2.10 / $0.21 / $6.60, a faster and pricier sibling to the existing GLM-5.2 line.
The same capture recorded a price cut on an existing model: GLM-5.2’s own cache-input rate fell from $0.26 to $0.14 per 1M tokens — a 46% reduction — while its $1.40 input and $4.40 output rates held steady. No other pricing surface moved: dedicated deployment GPU/CPU rates, the on-demand Training rate card, and the Basic/Pro/Enterprise tier structure are all unchanged from the prior capture.
Eight days after pruning its Model APIs catalog to eight models, Baseten added flagship Kimi K3 ($3.00 input / $0.30 cache input / $15.00 output per 1M tokens) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60), and cut GLM-5.2's own cache-input rate 46% to $0.14 per 1M tokens (from $0.26; its $1.40 input and $4.40 output rates held steady). Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.