Ask
Price change

Together AI runs a 27% promo on Dedicated Inference H100 through September 30

Together AI pricing

Together AI cut Dedicated Inference on-demand HGX H100 pricing from $5.49/hr to $3.99/hr (27% off) in a promotion running through September 30, 2026.

Before

Dedicated Inference on-demand HGX H100: $5.49/hr (list)

After

Dedicated Inference on-demand HGX H100: $3.99/hr promotional rate, with a "PROMO" badge and the note "PROMOTION VALID UNTIL 09/30/26" on the pricing page; list price $5.49/hr resumes after

Together’s Dedicated Inference table now shows a “PROMO” badge on the NVIDIA HGX H100 on-demand row, with the page’s own caption reading “PROMOTION VALID UNTIL 09/30/26”. The rate card displays a struck-through $5.49/hr next to a live $3.99/hr — the same figure Together already charges for on-demand H100 in its separate multi-node GPU Clusters product, effectively price-matching single-tenant dedicated inference to cluster rates for the promotion’s duration. No other hardware line (B200 dedicated, any reserved tenor, or any GPU Cluster rate) carries the discount, and the page gives no indication of what happens after October 1 beyond the struck-through list price shown alongside it.

The same capture cycle found the serverless Chat catalog rotating as usual — DeepSeek V4 Pro (original), Kimi K2.7 Code, NVIDIA Nemotron 3 Ultra, Inkling Small, Gemma-4-31B-it-Pearl, Gemma 3n E4B Instruct, and gpt-oss-20B dropped off (cross-confirmed against Together’s docs catalog), while GLM-5.3, GLM-5.3-Flash, and Qwen3.8 Flash joined — and one surviving model, Qwen3.8-2.4T-A95B, repriced from $2.50/$6.25 per 1M tokens to $2.00/$6.00, a 20% input cut.

From Together AI's pricing timeline
Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced

Together's Dedicated Inference table now tags on-demand NVIDIA HGX H100 with a "PROMO" badge reading "PROMOTION VALID UNTIL 09/30/26," showing a struck-through $5.49 list price next to a live $3.99/hr — a 27% cut with a published end date. On-demand B200 dedicated ($8.99/hr) and every reserved/Contact-us row are unaffected, and GPU Clusters, Provisioned Throughput, Code Sandbox, Storage, Image, Video, and Transcribe all confirmed byte-identical to 2026-08-26. The serverless Chat catalog rotated: DeepSeek V4 Pro (original), Kimi K2.7 Code, NVIDIA Nemotron 3 Ultra, Inkling Small (also dropped from Audio/speech), Gemma-4-31B-it-Pearl, Gemma 3n E4B Instruct, and gpt-oss-20B are confirmed absent from both the live page and the docs catalog; GLM-5.3, GLM-5.3-Flash, and Qwen3.8 Flash joined. One survivor was genuinely repriced — Qwen3.8-2.4T-A95B fell from $2.50 input / $6.25 output ($0.50 cached) to $2.00 input / $6.00 output ($0.25 cached), a 20% input cut, though the docs catalog still shows the old $0.50 cached-input rate (a new docs-vs-live conflict on that figure only). The Specialized fine-tuning tab's selector initially failed to actuate again (an exact-case mismatch on the tab's accessible name); a case-insensitive selector fix re-verified it later the same day, with every specialized-tier rate confirmed byte-identical to 2026-08-26.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity