Together AI runs a 27% promo on Dedicated Inference H100 through September 30
Together AI cut Dedicated Inference on-demand HGX H100 pricing from $5.49/hr to $3.99/hr (27% off) in a promotion running through September 30, 2026.
Dedicated Inference on-demand HGX H100: $5.49/hr (list)
Dedicated Inference on-demand HGX H100: $3.99/hr promotional rate, with a "PROMO" badge and the note "PROMOTION VALID UNTIL 09/30/26" on the pricing page; list price $5.49/hr resumes after
Together’s Dedicated Inference table now shows a “PROMO” badge on the NVIDIA HGX H100 on-demand row, with the page’s own caption reading “PROMOTION VALID UNTIL 09/30/26”. The rate card displays a struck-through $5.49/hr next to a live $3.99/hr — the same figure Together already charges for on-demand H100 in its separate multi-node GPU Clusters product, effectively price-matching single-tenant dedicated inference to cluster rates for the promotion’s duration. No other hardware line (B200 dedicated, any reserved tenor, or any GPU Cluster rate) carries the discount, and the page gives no indication of what happens after October 1 beyond the struck-through list price shown alongside it.
The same capture cycle found the serverless Chat catalog rotating as usual — DeepSeek V4 Pro (original), Kimi K2.7 Code, NVIDIA Nemotron 3 Ultra, Inkling Small, Gemma-4-31B-it-Pearl, Gemma 3n E4B Instruct, and gpt-oss-20B dropped off (cross-confirmed against Together’s docs catalog), while GLM-5.3, GLM-5.3-Flash, and Qwen3.8 Flash joined — and one surviving model, Qwen3.8-2.4T-A95B, repriced from $2.50/$6.25 per 1M tokens to $2.00/$6.00, a 20% input cut.
Together's Dedicated Inference table now tags on-demand NVIDIA HGX H100 with a "PROMO" badge reading "PROMOTION VALID UNTIL 09/30/26," showing a struck-through $5.49 list price next to a live $3.99/hr — a 27% cut with a published end date. On-demand B200 dedicated ($8.99/hr) and every reserved/Contact-us row are unaffected, and GPU Clusters, Provisioned Throughput, Code Sandbox, Storage, Image, Video, and Transcribe all confirmed byte-identical to 2026-08-26. The serverless Chat catalog rotated: DeepSeek V4 Pro (original), Kimi K2.7 Code, NVIDIA Nemotron 3 Ultra, Inkling Small (also dropped from Audio/speech), Gemma-4-31B-it-Pearl, Gemma 3n E4B Instruct, and gpt-oss-20B are confirmed absent from both the live page and the docs catalog; GLM-5.3, GLM-5.3-Flash, and Qwen3.8 Flash joined. One survivor was genuinely repriced — Qwen3.8-2.4T-A95B fell from $2.50 input / $6.25 output ($0.50 cached) to $2.00 input / $6.00 output ($0.25 cached), a 20% input cut, though the docs catalog still shows the old $0.50 cached-input rate (a new docs-vs-live conflict on that figure only). The Specialized fine-tuning tab's selector initially failed to actuate again (an exact-case mismatch on the tab's accessible name); a case-insensitive selector fix re-verified it later the same day, with every specialized-tier rate confirmed byte-identical to 2026-08-26.