Ask
Price change

Together AI raises GPU cluster reserved rates, unifies speech-to-text pricing

Together AI pricing

Together AI raised GPU cluster reserved H100 rates and unified speech-to-text pricing at $0.0015/audio-minute, cutting Nemotron 3.5 ASR pricing 67%.

Before

GPU Cluster reserved H100: $3.59/hr (7-30d), $3.29 (31-90d), $3.09 (91-180d); STT metered per-character for Parakeet ($0.0035/1M chars) and Nemotron 3 ASR ($0.0015/1M chars), per-minute for Whisper ($0.0015/min), Whisper Streaming ($0.0035/min), and Nemotron 3.5 ASR ($0.0045/min)

After

GPU Cluster reserved H100: $3.69/hr (7-30d), $3.45 (31-90d), $3.19 (91-180d); all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, Nemotron 3 ASR Streaming, Nemotron 3.5 ASR Streaming) now bill at a unified $0.0015 per audio minute

Proof of change

Together AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 21, 2026 Captured Jul 29, 2026 · 8 days apart
main-chat
$0.08 $0.07 $0.00 $0.02 $65.00 $15 $773,070 $1,802,370 $3.69 $3.45 $3.19
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 3 other pages we compared
main-finetuning-specialized
$0.08 $0.07 $0.00 $0.02 $65.00 $15 $773,070 $1,802,370 $3.69 $3.45 $3.19
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

main-image
$1.74 $0.20 $3.48 $0.30 $0.95 $0.19 $773,070 $1,802,370 $3.69 $3.45 $3.19
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

models-docs
Captured Jul 21, 2026
Captured Jul 29, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Together’s managed GPU Cluster reserved H100 rate rose for the first time since the aggressive back-to-back cuts of June 2026 — up to $3.69/hr on a 7-30 day reservation (from $3.59), $3.45/hr on 31-90 days (from $3.29), and a $3.19/hr floor on 91-180 days (from $3.09). On-demand H100, and every published H200/B200 on-demand and reserved rate, held flat, so the increase is isolated to the reserved-H100 tenors.

In the same capture cycle, Together’s speech-to-text catalog was re-metered onto a single unit. Previously the four serverless STT models split across two different meters — per-1M-characters for Parakeet TDT 0.6B and Nemotron 3 ASR, per-audio-minute for Whisper and Nemotron 3.5 ASR — at four different rates. All four now bill at a flat $0.0015 per audio minute, which is a like-for-like 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a unit change (character to minute) for Parakeet and Nemotron 3 ASR. The serverless chat catalog also grew (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) without moving any previously-published per-model rate.

From Together AI's pricing timeline
GPU Cluster reserved H100 rates rise (first increase since June cuts) + speech-to-text unified to $0.0015/audio-minute

Together's GPU Cluster reserved H100 rate rose across all three commitment tenors for the first time since the back-to-back June 2026 cuts: 7–30 day $3.59→$3.69/hr, 31–90 day $3.29→$3.45/hr, and the 91–180 day floor $3.09→$3.19/hr. On-demand H100 and every H200/B200 on-demand and reserved rate held flat, so the increase is isolated to reserved-H100 tenors. In the same capture cycle, all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) were unified onto a single $0.0015-per-audio-minute meter — a 67% cut on Nemotron 3.5 ASR (from $0.0045/min) and a billing-unit change (from per-1M-characters to per-audio-minute) for Parakeet TDT 0.6B and Nemotron 3 ASR. Four new serverless chat models (Kimi K3, Gemma 4 31B, Qwen3.7-Plus, LFM2.5-8B-A1B) joined the catalog without moving any previously-published rate.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity