Databricks Foundation Model APIs add GLM-5.2, DeepSeek V4, Kimi K3/K2.7, and Qwen 3.5 122B; Model Serving adds H100 GPUs
Databricks (Mosaic AI) pricing
Databricks expanded Mosaic AI's Foundation Model APIs catalog with GLM-5.2, DeepSeek V4 Pro/Flash, Kimi K3/K2.7, and Qwen 3.5 122B, and added H100 GPU configurations to Model Serving — all still billed at the existing $0.07/DBU rate.
Foundation Model APIs catalog centered on Llama 4 Maverick, Llama 3.3 70B, GPT OSS 120B/20B, Qwen 3 Next 80B, Gemma 3 12B, Llama 3.1 8B, and embedding models (GTE, BGE Large, Qwen 3 0.6B); Model Serving GPU options topped out at A100 80GB x8 (628 DBU/hr).
Catalog now also includes Kimi K3 (42.857/214.286 DBU per 1M in/out), Kimi K2.7, GLM-5.2 (20.000/62.857) and a Priority variant, Inkling, DeepSeek V4 Pro (18.857/56.571) and V4 Flash, and Qwen 3.5 122B (3.143/31.429) plus a Priority variant; Model Serving now offers L40S, A10G x4, A100 40GB x8, and H100 x1/x8 (up to 800 DBU/hr) GPU configurations. The underlying $0.07/DBU AI rate and $0.65/DBU training rate are unchanged.
Databricks did not change its core DBU rate card, but it materially widened what that rate buys. The Foundation Model Serving page now lists roughly twice as many models as it did in June, adding several of 2026’s newer open-weight releases — GLM-5.2, DeepSeek V4 Pro/Flash, Kimi K3/K2.7, and Qwen 3.5 122B — alongside “Priority” pay-per-token variants of GLM-5.2 and Qwen 3.5 122B that carry higher DBU/1M-token rates for lower-latency serving. On the compute side, Model Serving’s GPU rate card grew from four listed configurations to nine, adding L40S x1, A10G x4, A100 40GB x8, and H100 x1/x8 (the top end now runs 800 DBU/hour, versus 628 DBU/hour for the previous ceiling of A100 80GB x8). Vector Search (AI Search) also now surfaces an explicit $/hour compute + $/GB/month storage breakdown (Standard: $0.28/hr + $0.230/GB/mo with the first 30GB free; Storage Optimized: $1.28/hr + $0.046/GB/mo) rather than only the DBU/hour figure. None of the existing model or GPU prices moved — this is catalog and hardware expansion, not a repricing.
Databricks added GLM-5.2, DeepSeek V4 Pro/Flash, Kimi K3/K2.7, and Qwen 3.5 122B to Foundation Model APIs, extended Model Serving's GPU rate card with L40S, A10G x4, A100 40GB x8, and H100 x1/x8 (ceiling rises from 628 to 800 DBU/hr), and gave Vector Search an explicit $/hour compute plus $/GB-month storage split. The core $0.07/DBU AI rate and $0.65/DBU training rate are unchanged — catalog and hardware expansion, not repricing.