Ask
Price change

Together AI cuts Kokoro-82M TTS 60% and adds two chat models

Together AI pricing

Together AI cut Kokoro-82M TTS from $10.00 to $4.00 per 1M characters and added Inkling ($1.00 in / $4.05 out) plus a 0.00-input PrismML model to its serverless card.

Before

Kokoro-82M TTS $10.00 per 1M characters; no Inkling or PrismML Ternary Bonsai 27B on the serverless rate card

After

Kokoro-82M TTS $4.00 per 1M characters; Inkling at $1.00 in / $4.05 out / $0.17 cached (and $1.00 per 1M characters for speech); PrismML Ternary Bonsai 27B listed at 0.00 input with no published output rate

Proof of change

Together AI's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 14, 2026 Captured Jul 21, 2026 · 7 days apart
main-chat
$1.00 $4.05
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 3 other pages we compared
main-finetuning-specialized
$1.00 $4.05
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

main-image
$1.00 $4.05
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

models-docs
$1.00 $4.05
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

The move is confined to the model catalog: Together’s compute rate card was unchanged week-over-week, with dedicated inference (on-demand HGX H100 $5.49/hr, HGX B200 $8.99/hr), GPU clusters (on-demand H100 $3.99/hr, reserved 7-30 day H100 $3.59/hr), Provisioned Throughput ($0.05/PTU-min on MiniMax M3 and GLM-5.2), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), storage ($0.16/GiB-month), and both fine-tuning tiers all holding their 2026-07-14 rates.

The 60% cut lands on the cheap end of the text-to-speech card - Kokoro-82M is now $4.00 per 1M characters against Orpheus TTS at $15 and Cartesia Sonic-2/Sonic-3 at $65.00, both unchanged. That widens the spread between the open-weight TTS line and the licensed commercial voices from roughly 6.5x to 16x, which reads as Together pricing its self-hosted-class speech models toward the marginal cost of serving them rather than toward the commercial voice market.

Alongside it, two models joined the chat rate card. Inkling is a conventional paid line ($1.00 input / $4.05 output per 1M tokens, with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech entry). PrismML Ternary Bonsai 27B is the unusual one: it is published at an input rate of 0.00 with no output rate shown at all - a free-to-call serving line on a rate card that otherwise prices every model.

From Together AI's pricing timeline
Kokoro-82M TTS cut 60% + two chat models added; compute rate card flat

Together's compute meters held flat week-over-week — dedicated HGX H100 at $5.49/hr and HGX B200 at $8.99/hr, GPU clusters at $3.99/hr on-demand H100 and $3.59/hr reserved, PTU at $0.05/PTU-min, Code Sandbox, storage, and both fine-tuning tiers all unchanged. The movement was entirely in the model catalog: Kokoro-82M TTS was cut 60% from $10.00 to $4.00 per 1M characters (against Orpheus at $15 and Cartesia Sonic-2/-3 unchanged at $65.00, stretching the TTS spread to ~16×), Inkling was added at $1.00 input / $4.05 output per 1M tokens with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech line, and PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with no output rate published.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro 0813 at $1.32/$3.96 with $0.13 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.2 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute — all four serverless speech-to-text models (Whisper Large v3, Parakeet TDT 0.6B v3, and both Nemotron 3 and 3.5 ASR Streaming variants) unified onto one flat $0.0015/min rate on 29 July 2026 — and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Sep 2026
    Dedicated Inference on-demand H100 gets a 27% promo through 2026-09-30; chat catalog rotates and Qwen3.8-2.4T-A95B repriced
  2. Aug 2026
    Compute rate card and full model catalog flat again; Specialized fine-tuning tab re-actuated after a two-week capture gap and shows catalog rotation, not repricing
  3. Aug 2026
    Compute rate card flat again; DeepSeek V4 Pro 0813 + 6 HappyHorse video models added; GLM-5.1 and Kimi K2.6 rotated off the catalog
  4. Aug 2026
    Compute rate card flat a fifth straight capture; Cogito v2.1 671B, Rnj-1 Instruct, and Inkling Small's dual-meter billing confirmed
  5. Aug 2026
    Compute rate card flat a fourth straight week; Qwen3.8-2.4T-A95B and Inkling Small added
Full Together AI timeline

More Together AI activity

All pricing activity