Ask
Launch

Together AI launches Provisioned Throughput (PTU) and cuts Dedicated Inference rates

Together AI pricing

Together AI launched a Provisioned Throughput (PTU) SKU at $0.05/PTU-min and cut Dedicated Inference rates — on-demand H100 to $5.49/hr, B200 to $8.99/hr.

Before

No reserved-throughput product; Dedicated Inference priced per instance — 1x H100 80GB $6.49/hr, 1x HGX B200 180GB $11.95/hr, H200 Contact us.

After

New Provisioned Throughput SKU reserves capacity in throughput units at $0.05/PTU-minute (MiniMax M3, GLM-5.2); Dedicated Inference restructured to per-GPU-per-hour, on-demand vs reserved — on-demand H100 $5.49/hr, B200 $8.99/hr, with H200/B300/GB200/GB300 quoted Contact us and all reserved capacity Contact sales.

Proof of change

Together AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jun 30, 2026 Captured Jul 14, 2026 · 14 days apart
models-docs
Captured Jun 30, 2026
Captured Jul 14, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 3 other pages we compared
main-chat
Captured Jun 30, 2026
Captured Jul 14, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

main-finetuning-specialized
Captured Jun 30, 2026
Captured Jul 14, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

main-image
Captured Jun 30, 2026
Captured Jul 14, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Together AI added a fifth pricing surface, Provisioned Throughput (PTU), which reserves dedicated inference capacity in throughput units billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2). Each PTU delivers a fixed, model-specific tokens-per-minute rate, and an on-page calculator sizes the PTUs required for a traffic profile and estimates monthly cost and savings versus a commercial model’s list price (assuming 24/7 provisioning).

In the same update, Dedicated Inference was restructured from a per-instance table into a per-GPU-per-hour grid split into on-demand (pay-as-you-go) and reserved (Contact sales) columns. On-demand HGX H100 dropped from $6.49 to $5.49/hr (-15%) and HGX B200 from $11.95 to $8.99/hr (-25%), while newly listed HGX H200, HGX B300, GB200 NVL72, and GB300 NVL72 lines are quoted “Contact us”. Serverless per-token and GPU Cluster rates were unchanged. The pricing page also carries a new Series C funding banner.

From Together AI's pricing timeline
Provisioned Throughput (PTU) launch + Dedicated Inference restructure & price cuts

Together launched Provisioned Throughput — a new SKU that reserves dedicated capacity in throughput units (PTUs) billed per PTU-minute ($0.05/PTU-min on MiniMax M3 and GLM-5.2), with an on-page calculator that sizes PTUs and estimates monthly cost vs. commercial-model list prices. In the same update, Dedicated Inference was restructured to a per-GPU-per-hour table split into on-demand (pay-as-you-go) vs reserved (Contact sales) columns: on-demand HGX H100 fell to $5.49/hr (from $6.49) and HGX B200 to $8.99/hr (from $11.95), and NVIDIA HGX H200, HGX B300, GB200 NVL72, and GB300 NVL72 lines were added (quoted "Contact us"). GPU Cluster and serverless rates were unchanged. The pricing page also carries a new Series C funding banner.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity