Together AI reprices serverless and cuts GPU cluster rates
Together cut GPU cluster rates (on-demand H100 $5.49→$4.79, reserved 7–30d H100 $4.99→$4.19), repriced serverless models, and began publishing cached-input rates on serverless inference.
On-demand H100 $5.49/hr, reserved 7–30d H100 $4.99/hr, B200 reserved $9.65/hr; DeepSeek V4 Pro $2.10/$4.40; Llama 3.3 70B $0.88/$0.88; no published cached-input discount.
On-demand H100 $4.79/hr, reserved 7–30d H100 $4.19/hr (as low as $3.29/hr on 91–180d), B200 reserved $7.99/hr; DeepSeek V4 Pro $1.74/$3.48 with $0.20 cached input; Llama 3.3 70B $1.04/$1.04; cached-input rates now published.
Proof of change
Together AI's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 1 other page we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
- 3 pages had no counterpart in the earlier capture and are not shown.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Together AI repriced its serverless rate card and cut GPU cluster rates across the board. On-demand cluster H100 dropped to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 reaching $3.29/hr on a 91–180 day reservation. H200 now appears on the cluster rate card (on-demand $5.99/hr, reserved $4.99–$3.99/hr).
On serverless, DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40), Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15), and Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). Most notably, Together now publishes cached-input rates on serverless models (e.g. DeepSeek V4 Pro $0.20, GLM-5.1/5.2 $0.26, Kimi K2.6 $0.20) — closing the previously-flagged competitive gap against Fireworks, OpenAI, and Anthropic, which all shipped cached-input discounts earlier.
Dedicated endpoint rates (H100 $6.49/hr, HGX B200 180GB $11.95/hr), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), Code Interpreter ($0.03/session), storage ($0.16/GiB-month), and the standard fine-tuning rate card were unchanged.
Together repriced its serverless rate card and cut GPU cluster rates. DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40) and now shows a $0.20 cached-input rate; Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15); Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). On-demand cluster H100 fell to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 as low as $3.29/hr on a 91–180 day reservation. Cached-input pricing is now published on serverless models.