Together AI cuts Kokoro-82M TTS 60% and adds two chat models
Together AI cut Kokoro-82M TTS from $10.00 to $4.00 per 1M characters and added Inkling ($1.00 in / $4.05 out) plus a 0.00-input PrismML model to its serverless card.
Kokoro-82M TTS $10.00 per 1M characters; no Inkling or PrismML Ternary Bonsai 27B on the serverless rate card
Kokoro-82M TTS $4.00 per 1M characters; Inkling at $1.00 in / $4.05 out / $0.17 cached (and $1.00 per 1M characters for speech); PrismML Ternary Bonsai 27B listed at 0.00 input with no published output rate
Proof of change
Together AI's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
Show 3 other pages we compared
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Showing the whole page as captured — scroll either panel, or open it at full size.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
The move is confined to the model catalog: Together’s compute rate card was unchanged week-over-week, with dedicated inference (on-demand HGX H100 $5.49/hr, HGX B200 $8.99/hr), GPU clusters (on-demand H100 $3.99/hr, reserved 7-30 day H100 $3.59/hr), Provisioned Throughput ($0.05/PTU-min on MiniMax M3 and GLM-5.2), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), storage ($0.16/GiB-month), and both fine-tuning tiers all holding their 2026-07-14 rates.
The 60% cut lands on the cheap end of the text-to-speech card - Kokoro-82M is now $4.00 per 1M characters against Orpheus TTS at $15 and Cartesia Sonic-2/Sonic-3 at $65.00, both unchanged. That widens the spread between the open-weight TTS line and the licensed commercial voices from roughly 6.5x to 16x, which reads as Together pricing its self-hosted-class speech models toward the marginal cost of serving them rather than toward the commercial voice market.
Alongside it, two models joined the chat rate card. Inkling is a conventional paid line ($1.00 input / $4.05 output per 1M tokens, with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech entry). PrismML Ternary Bonsai 27B is the unusual one: it is published at an input rate of 0.00 with no output rate shown at all - a free-to-call serving line on a rate card that otherwise prices every model.
Together's compute meters held flat week-over-week — dedicated HGX H100 at $5.49/hr and HGX B200 at $8.99/hr, GPU clusters at $3.99/hr on-demand H100 and $3.59/hr reserved, PTU at $0.05/PTU-min, Code Sandbox, storage, and both fine-tuning tiers all unchanged. The movement was entirely in the model catalog: Kokoro-82M TTS was cut 60% from $10.00 to $4.00 per 1M characters (against Orpheus at $15 and Cartesia Sonic-2/-3 unchanged at $65.00, stretching the TTS spread to ~16×), Inkling was added at $1.00 input / $4.05 output per 1M tokens with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech line, and PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with no output rate published.