Price change

Together AI cuts Kokoro-82M TTS 60% and adds two chat models

Together AI pricing

Together AI cut Kokoro-82M TTS from $10.00 to $4.00 per 1M characters and added Inkling ($1.00 in / $4.05 out) plus a 0.00-input PrismML model to its serverless card.

Before

Kokoro-82M TTS $10.00 per 1M characters; no Inkling or PrismML Ternary Bonsai 27B on the serverless rate card

After

Kokoro-82M TTS $4.00 per 1M characters; Inkling at $1.00 in / $4.05 out / $0.17 cached (and $1.00 per 1M characters for speech); PrismML Ternary Bonsai 27B listed at 0.00 input with no published output rate

The move is confined to the model catalog: Together’s compute rate card was unchanged week-over-week, with dedicated inference (on-demand HGX H100 $5.49/hr, HGX B200 $8.99/hr), GPU clusters (on-demand H100 $3.99/hr, reserved 7-30 day H100 $3.59/hr), Provisioned Throughput ($0.05/PTU-min on MiniMax M3 and GLM-5.2), Code Sandbox ($0.0446/vCPU-hour, $0.0149/GiB-hour), storage ($0.16/GiB-month), and both fine-tuning tiers all holding their 2026-07-14 rates.

The 60% cut lands on the cheap end of the text-to-speech card - Kokoro-82M is now $4.00 per 1M characters against Orpheus TTS at $15 and Cartesia Sonic-2/Sonic-3 at $65.00, both unchanged. That widens the spread between the open-weight TTS line and the licensed commercial voices from roughly 6.5x to 16x, which reads as Together pricing its self-hosted-class speech models toward the marginal cost of serving them rather than toward the commercial voice market.

Alongside it, two models joined the chat rate card. Inkling is a conventional paid line ($1.00 input / $4.05 output per 1M tokens, with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech entry). PrismML Ternary Bonsai 27B is the unusual one: it is published at an input rate of 0.00 with no output rate shown at all - a free-to-call serving line on a rate card that otherwise prices every model.

From Together AI's pricing timeline
Kokoro-82M TTS cut 60% + two chat models added; compute rate card flat

Together's compute meters held flat week-over-week — dedicated HGX H100 at $5.49/hr and HGX B200 at $8.99/hr, GPU clusters at $3.99/hr on-demand H100 and $3.59/hr reserved, PTU at $0.05/PTU-min, Code Sandbox, storage, and both fine-tuning tiers all unchanged. The movement was entirely in the model catalog: Kokoro-82M TTS was cut 60% from $10.00 to $4.00 per 1M characters (against Orpheus at $15 and Cartesia Sonic-2/-3 unchanged at $65.00, stretching the TTS spread to ~16×), Inkling was added at $1.00 input / $4.05 output per 1M tokens with a $0.17 cached-input rate and a parallel $1.00 per-1M-character speech line, and PrismML Ternary Bonsai 27B was listed at an input rate of 0.00 with no output rate published.

About Together AI
together.ai ↗

Together AI runs a multi-SKU pure-usage cloud: per-token serverless inference for popular open-weight models (Llama 3.3 70B at $1.04/$1.04, DeepSeek V4 Pro at $1.74/$3.48 with $0.20 cached input, Qwen3.5 9B at $0.17/$0.25, GLM-5.1 at $1.40/$4.40), image generation billed per image or per megapixel (FLUX.2 [dev] $0.0154 per image, FLUX.1 [schnell] $0.0027 per megapixel, SD XL $0.0019 per megapixel), and per-hour dedicated and cluster GPUs. The same serverless API meters speech per 1M characters (Cartesia Sonic-2 and Sonic-3 at $65.00, Orpheus TTS at $15, Kokoro-82M at $4.00 after a 60% cut on 21 July 2026 — a roughly 16x spread that makes model choice, not volume, the driver of a speech bill), transcription per audio minute (Whisper Large v3 at $0.0015/min), and video per video (Google Veo 3.0 at $1.60, Sora 2 at $0.80).

Free tier
Yes
Commits
Available
Transparency
public

Together AI pricing history

  1. Jul 2026
    Kokoro-82M TTS cut 60% + two chat models added; compute rate card flat
  2. Jul 2026
    Provisioned Throughput (PTU) launch + Dedicated Inference restructure & price cuts
  3. Jul 2026
    $800M raise at an $8.3B valuation (IPO not ruled out)
  4. Jun 2026
    Reserved + on-demand GPU cluster rate cut (H100 down 12–16%)
  5. Jun 2026
    Serverless re-pricing + GPU cluster rate cuts + cached input
Full Together AI timeline

More Together AI activity

All pricing activity