Ask
Price change

Together AI reprices serverless and cuts GPU cluster rates

Together AI pricing

Together cut GPU cluster rates (on-demand H100 $5.49→$4.79, reserved 7–30d H100 $4.99→$4.19), repriced serverless models, and began publishing cached-input rates on serverless inference.

Before

On-demand H100 $5.49/hr, reserved 7–30d H100 $4.99/hr, B200 reserved $9.65/hr; DeepSeek V4 Pro $2.10/$4.40; Llama 3.3 70B $0.88/$0.88; no published cached-input discount.

After

On-demand H100 $4.79/hr, reserved 7–30d H100 $4.19/hr (as low as $3.29/hr on 91–180d), B200 reserved $7.99/hr; DeepSeek V4 Pro $1.74/$3.48 with $0.20 cached input; Llama 3.3 70B $1.04/$1.04; cached-input rates now published.

Proof of change

Together AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured May 29, 2026 Captured Jun 24, 2026 · 26 days apart
main
$2.10 $1.00 $0.10 $2.00 $0.88 $5.49 $1.74 $3.48 $0.28 $0.86 $0.95 $0.19
Captured May 29, 2026
Captured Jun 24, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 1 other page we compared
models-docs
$7.50 $1.00 $2.10 $0.88 $3.75 $0.17 $0.25 $0.95 $0.26 $1.74
Captured May 29, 2026
Captured Jun 24, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • 3 pages had no counterpart in the earlier capture and are not shown.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Together AI repriced its serverless rate card and cut GPU cluster rates across the board. On-demand cluster H100 dropped to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 reaching $3.29/hr on a 91–180 day reservation. H200 now appears on the cluster rate card (on-demand $5.99/hr, reserved $4.99–$3.99/hr).

On serverless, DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40), Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15), and Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). Most notably, Together now publishes cached-input rates on serverless models (e.g. DeepSeek V4 Pro $0.20, GLM-5.1/5.2 $0.26, Kimi K2.6 $0.20) — closing the previously-flagged competitive gap against Fireworks, OpenAI, and Anthropic, which all shipped cached-input discounts earlier.

Dedicated endpoint rates (H100 $6.49/hr, HGX B200 180GB $11.95/hr), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), Code Interpreter ($0.03/session), storage ($0.16/GiB-month), and the standard fine-tuning rate card were unchanged.

From Together AI's pricing timeline
Serverless re-pricing + GPU cluster rate cuts + cached input

Together repriced its serverless rate card and cut GPU cluster rates. DeepSeek V4 Pro dropped to $1.74/$3.48 (from $2.10/$4.40) and now shows a $0.20 cached-input rate; Qwen3.5 9B rose to $0.17/$0.25 (from $0.10/$0.15); Llama 3.3 70B rose to $1.04/$1.04 (from $0.88/$0.88). On-demand cluster H100 fell to $4.79/hr (from $5.49) and B200 to $8.19/hr (from $9.95); reserved 7–30 day H100 fell to $4.19/hr (from $4.99) and B200 to $7.99/hr (from $9.65), with H100 as low as $3.29/hr on a 91–180 day reservation. Cached-input pricing is now published on serverless models.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity