Baseten adds Kimi K3 and GLM-5.2 Fast, cuts GLM-5.2 cache pricing 46%
Baseten grew its Model APIs catalog to ten models with new flagship Kimi K3 and GLM-5.2 Fast, and cut GLM-5.2's cache-input rate 46% to $0.14 per 1M tokens.
8 Model APIs SKUs: Inkling, GLM-5.2 ($1.40 in / $0.26 cache / $4.40 out), GLM 4.7, Kimi K2.7 Code, Kimi K2.6, NVIDIA Nemotron 3 Ultra, DeepSeek V4, GPT OSS 120B.
10 Model APIs SKUs adding Kimi K3 ($3.00 in / $0.30 cache / $15.00 out) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60); GLM-5.2 cache input cut to $0.14 (input/output unchanged at $1.40/$4.40).
Eight days after pruning its Model APIs catalog from eleven models to eight, Baseten expanded it again — this time by addition rather than retirement. Kimi K3 launched as the platform’s new flagship model at $3.00 input / $0.30 cache input / $15.00 output per 1M tokens, flagged by a homepage “Kimi K3 is here.” banner, and became the most expensive SKU on the rate card by a wide margin (72% above the prior ceiling, DeepSeek V4, on input; over 3x on output). GLM-5.2 Fast joined alongside it at $2.10 / $0.21 / $6.60, a faster and pricier sibling to the existing GLM-5.2 line.
The same capture recorded a price cut on an existing model: GLM-5.2’s own cache-input rate fell from $0.26 to $0.14 per 1M tokens — a 46% reduction — while its $1.40 input and $4.40 output rates held steady. No other pricing surface moved: dedicated deployment GPU/CPU rates, the on-demand Training rate card, and the Basic/Pro/Enterprise tier structure are all unchanged from the prior capture.
Eight days after pruning its Model APIs catalog to eight models, Baseten added flagship Kimi K3 ($3.00 input / $0.30 cache input / $15.00 output per 1M tokens) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60), and cut GLM-5.2's own cache-input rate 46% to $0.14 per 1M tokens (from $0.26; its $1.40 input and $4.40 output rates held steady). Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.