Together AI raises GPU cluster reserved rates, unifies speech-to-text pricing
Together AI raised GPU cluster reserved H100 rates and unified speech-to-text pricing at $0.0015/audio-minute, cutting Nemotron 3.5 ASR pricing 67%.
GPU Cluster reserved H100: $3.59/hr (7-30d), $3.29 (31-90d), $3.09 (91-180d); STT metered per-character for Parakeet ($0.0035/1M chars) and Nemotron 3 ASR ($0.0015/1M chars), per-minute for Whisper ($0.0015/min), Whisper Streaming ($0.0035/min), and Nemotron 3.5 ASR ($0.0045/min)
GPU Cluster reserved H100: $3.69/hr (7-30d), $3.45 (31-90d), $3.19 (91-180d); all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, Nemotron 3 ASR Streaming, Nemotron 3.5 ASR Streaming) now bill at a unified $0.0015 per audio minute
Proof of change
Together AI's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 3 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Together’s managed GPU Cluster reserved H100 rate rose for the first time since the aggressive back-to-back cuts of June 2026 — up to $3.69/hr on a 7-30 day reservation (from $3.59), $3.45/hr on 31-90 days (from $3.29), and a $3.19/hr floor on 91-180 days (from $3.09). On-demand H100, and every published H200/B200 on-demand and reserved rate, held flat, so the increase is isolated to the reserved-H100 tenors.
In the same capture cycle, Together’s speech-to-text catalog was re-metered onto a single unit. Previously the four serverless STT models split across two different meters — per-1M-characters for Parakeet TDT 0.6B and Nemotron 3 ASR, per-audio-minute for Whisper and Nemotron 3.5 ASR — at four different rates. All four now bill at a flat $0.0015 per audio minute, which is a like-for-like 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a unit change (character to minute) for Parakeet and Nemotron 3 ASR. The serverless chat catalog also grew (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) without moving any previously-published per-model rate.
Together's GPU Cluster reserved H100 rate rose across all three commitment tenors for the first time since the back-to-back June 2026 cuts: 7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, and the 91–180 day floor $3.09→$3.19/hr. On-demand H100 and every H200/B200 on-demand and reserved rate held flat, so the increase is isolated to reserved-H100 tenors. In the same capture cycle, all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a single $0.0015-per-audio-minute meter — a 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a billing-unit change (from per-1M-characters to per-audio-minute) for Parakeet TDT 0.6B and Nemotron 3 ASR. Four new serverless chat models (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) joined the catalog without moving any previously-published rate.