Ask
Price change

Together AI boosts Provisioned Throughput capacity, adds Kimi K3

Together AI pricing

Together AI added Kimi K3 as a third Provisioned Throughput model and raised per-PTU capacity 18-80% for MiniMax M3 and GLM-5.2 while holding the $0.05/PTU-minute sticker price flat.

Before

Provisioned Throughput covered MiniMax M3 (138,840 input / 694,200 cached / 23,140 output TPM/PTU) and GLM-5.2 (35,731 / 192,400 / 9,620 TPM/PTU), both at $0.05/PTU-minute

After

Provisioned Throughput covers MiniMax M3 (166,667 / 833,333 / 41,667 TPM/PTU, +20-80%), GLM-5.2 (35,714 / 192,308 / 11,364 TPM/PTU, output +18%), and new third model Kimi K3 (16,667 / 166,667 / 3,333 TPM/PTU) — all still at $0.05/PTU-minute

Proof of change

Together AI's pricing pages, as we captured them on two dates.

Capture only
Captured Aug 11, 2026 Captured Aug 12, 2026 · 1 days apart
models-docs
$0.35 $1.50 $0.11

These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.

Captured Aug 11, 2026
Captured Aug 12, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 1 other page we compared
main-chat
Captured Aug 11, 2026
Captured Aug 12, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • These prices come from our capture alone — they were not confirmed against an independent second source.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Together’s Provisioned Throughput SKU — launched in July 2026 as a fixed tokens-per-minute reservation billed per PTU-minute — gained its third supported model on 2026-08-12: Kimi K3 joins MiniMax M3 and GLM-5.2 on the on-page PTU sizing calculator and its underlying compute-costs reference table, all still priced at $0.05 per PTU-minute.

The more consequential move is a capacity increase on the two existing models. MiniMax M3’s published per-PTU capacity rose from 138,840 to 166,667 input tokens-per-minute and from 694,200 to 833,333 cached tokens-per-minute (both +20%), and from 23,140 to 41,667 output tokens-per-minute (+80%). GLM-5.2’s output capacity rose from 9,620 to 11,364 tokens-per-minute (+18%), with input and cached capacity effectively unchanged. Because the $0.05/PTU-minute sticker price did not move, a buyer reserving the same number of PTUs now gets meaningfully more throughput for the same bill — an effective price cut on the unit economics of reserved capacity, even though the headline rate card shows no change. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week, and the chat/image/video serverless catalogs continued growing without moving any other previously-published rate.

From Together AI's pricing timeline
Provisioned Throughput capacity increase + Kimi K3 added; PrismML confirmed Free

Together added Kimi K3 as a third Provisioned Throughput model and increased published per-PTU capacity for MiniMax M3 (input 138,840→ 166,667 and cached 694,200→833,333 TPM/PTU, +20% each; output 23,140→ 41,667 TPM/PTU, +80%) and GLM-5.2 (output 9,620→11,364 TPM/PTU, +18%), while holding the $0.05/PTU-minute sticker price flat for all three models — an effective cut in cost per unit of reserved throughput. Together's model-catalog docs also newly confirmed PrismML Ternary Bonsai 27B as fully "Free" (input and output), resolving the ambiguous blank-output cell flagged since 2026-07-21. Every other compute meter (GPU Clusters, Dedicated Inference, Code Sandbox, Storage, both fine-tuning tiers) held its 2026-07-29 rate for a third consecutive week; the specialized fine-tuning tier gained three new per-model rows (Llama 4 Maverick, Qwen3-Coder-480B-A35B-Instruct, Qwen3.5-122B-A10B) and the chat/image/video catalogs continued growing without moving any other previously-published rate.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity