The voice minute is unbundling into pass-through component meters
Self-serve voice-agent platforms are decomposing the per-minute price into a thin platform fee plus at-cost pass-through of the LLM, TTS and telephony underneath — Vapi's advertised $0.05/min is only its hosting fee; real deployments run ~$0.13–$0.33/min. Bring-your-own components is a published discount lever, priced as exactly as Deepgram's $0.010/min BYO-TTS delta. A counter-camp (Cartesia, Bland, Tavus) still sells one bundled minute at one price.
What's happening — and why
What's happening: the per-minute price of a voice agent is splitting into a bill of materials. Retell itemizes a minute into Voice Infra ($0.055/min) plus separately priced LLM, TTS, telephony and add-on lines, metered to the nearest second; Vapi charges $0.05/min for hosting and passes model and telephony through at cost; Synthflow abolished its platform fee entirely and bills the Voice Engine, the LLM and telephony as three separate per-minute meters; Deepgram publishes managed and BYO-component tiers side by side, exactly $0.010/min apart; ElevenLabs keeps cutting its Agents pricing while steering new subscriptions to pay-as-you-go minutes.
Why: the components of a conversation minute — speech-to-text, the LLM turn, text-to-speech, the phone line — have wildly different and fast-moving costs, and many buyers already hold their own Twilio contract or model key. Pricing each component separately lets the platform shrink its own fee to a defensible orchestration margin, transmit component-price deflation automatically, and turn bring-your-own keys into a published, predictable discount instead of a negotiation — voice AI's version of the BYOK split between orchestration and inference. The counter-camp (Cartesia, Bland, Tavus) bets the other way: buyers want a quote, not a bill of materials, so they sell one flat bundled minute.
How it works
Evidence over time
11 supporting · 5 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| Synthflow | Jun 2025 | Killed cheap flat tiers (Starter was $29/mo) and moved to no-platform-fee pay-as-you-go that bills the Voice Engine, the LLM, and telephony as three separate per-minute meters; BYO Twilio zeroes the telephony line. |
| Retell AI | Jun 2026 | Pricing fully unbundled: a voice-agent minute is itemized into Retell Voice Infra ($0.055/min) plus separately-priced LLM, TTS, telephony and add-ons, metered to the nearest second. |
| Vapi | Jun 2026 | Headline $0.05/min is only the hosting fee — model and telephony costs are passed through at cost, so deployments run ~$0.13–$0.33/min depending on the stack assembled. |
| Deepgram | May 2026 | Voice Agent API sells managed vs BYO component tiers side by side — Standard – BYO TTS at $0.065/min vs $0.075/min fully managed — pricing the bundled component as an explicit per-minute delta. |
| ElevenLabs | May 2026 | Cut Conversational AI (Agents) pricing again while steering new subscriptions to pay-as-you-go per-minute agent usage — the platform-minute fee keeps thinning. |
| Bland AI | Jul 2026 | THE COUNTEREXAMPLE FLIPS — see counterexamples for the prior state. Bland dropped telephony out of its all-in per-minute rate: the rate now covers LLM + STT + TTS only, and the pricing-page FAQ states 'Telephony is billed separately, on your own carrier or Bland's at pass-through cost.' Previously telephony (PSTN and SIP) was explicitly listed under 'Everything included in your per-minute rate.' The bundled camp's clearest member adopted the pass-through mechanic on exactly the component every unbundler starts with. |
| LiveKit | Jul 2026 | The pass-through catalog expands and the advertised price becomes a calculator output. LiveKit Inference added six per-minute LLM SKUs — Kimi K2.6 $0.0035/min, GPT-5.6 Luna $0.0040, Grok 4.3 $0.0042, Grok 4.5 $0.0070, GPT-5.6 Terra $0.0101, GPT-5.6 Sol $0.0203 — and the pricing-page calculator's default total shifted from $0.0735/min (GPT-5.3 Chat stack) to $0.0672/min (Gemma 4 31B stack). Cloud plan fees unchanged: Build $0 / Ship $50 / Scale $500 / Enterprise custom, over a flat ~$0.01/min platform take. |
| LiveKit | Aug 2026 | THE PROOF THAT THE HEADLINE IS A BILL OF MATERIALS: an 80% component cut that did not move the advertised price. LiveKit cut GPT-5.6 Luna from $0.0040/min to $0.0008/min (-80%) and GPT-5.6 Terra from $0.0101 to $0.0081 (-20%), with GPT-5.6 Sol held at $0.0203 — a genuine per-SKU repricing, not the catalog-default swap it used in July to move its advertised blended rate. And the pricing page's default voice-agent estimate STILL READS $0.0672/min, because that calculation runs on Gemma 4 31B, which was not among the repriced models. No LiveKit Cloud plan price, allotment or overage rate moved: Build free, Ship $50, Scale $500, Enterprise custom, agent-session minutes $0.01. A vendor whose headline number fails to respond to an 80% cut in one of its components does not have a price; it has a parts list and a calculator. |
| LiveKit | Jul 2026 | A FREE component enters the pass-through catalog — the first $0-rated component in a voice bill of materials in this corpus. LiveKit Inference added Fish Audio TTS, including a free SKU, alongside two Gemini variants. When a component can be zero, the assembled floor collapses toward the platform take alone (LiveKit's is a flat ~$0.01/min), which is the logical endpoint of the thin-fee-plus-pass-through architecture and the hardest thing for a bundled per-minute rate to answer. |
| xAI | Aug 2026 | Unbundling gains a QUALITY dimension: the LLM component of a voice minute is now itself a ladder. xAI renamed its generic 'Realtime' line item to 'Speech to Speech' and split it across two explicitly named models — grok-voice-think-fast-1.0 unchanged at $0.05/minute ($3.00/hr audio) and a new grok-voice-think-fast-2.0 at $0.08/minute ($4.80/hr), a flat 60% premium — both still plus $0.004 per message of text input. Assembling a voice minute is no longer only a question of which vendor supplies each component; it is a question of which GRADE of that component. Sarvam ran the workflow version of the same move on 2026-08-15, doubling its Dubbing per-minute rate on an `editor_flow: true` request flag. |
| Speechmatics | Jul 2026 | Component-level deflation underneath the unbundled minute: cut real-time enhanced-accuracy STT, added a multilingual Batch Melia 1 model, and grew the free tier from 2,400 to 3,000 minutes per month (50 hours: 1,200 real-time + 1,800 batch). Every price cut at a component vendor lowers the floor of everyone else's assembled minute — and makes the bundled camp's single number look worse over time. |
Counterexamples
- Synthflow · Jun 2026 — THE LEAD EXAMPLE WITHDRAWS — but by exiting self-serve, not by re-bundling. The fully unbundled pay-as-you-go card (no platform fee; $0.09/min Voice Engine + LLM $0.02–$0.05/min + telephony $0.00–$0.02/min, ~$0.11–$0.24 all-in, 5 concurrent calls included) was replaced by a single Enterprise plan: contracts from $30,000/year, final pricing scoped per deployment (call volume, concurrency, telephony, integrations, security) via Contact Sales. No self-serve path and no published per-minute rate remain. The distinct failure mode: an unbundled minute is a configuration, and a configuration may need a salesperson.
- Cartesia · Jul 2026 — Bundling's repricing risk, demonstrated. Cartesia removed the Monthly/Yearly (Save 20%) toggle, so the self-serve panel went from defaulting to $4/$39/$239 yearly-equivalent rates to monthly-only Free $0 / Pro $5 / Startup $49 / Scale $299 / Enterprise custom. A ~20–25% effective increase for annual buyers, delivered as one opaque number with no component to interrogate — the exact hazard this trend attributes to the bundled camp.
- Cartesia · Feb 2026 — Took Voice Agents to GA on a flat, bundled per-minute rate — one number, no component meters.
- Bland AI · Dec 2025 — SUPERSEDED — PARTIALLY FLIPPED 2026-07-21, see evidence. As recorded here: moved from a single flat $0.09/min to tier-linked flat per-minute rates (free tier up 55% to $0.14/min) — repriced the bundle rather than unbundling it; compliance (HIPAA, SOC 2, PCI) stays included rather than itemized.
- Tavus · Feb 2026 — Deliberately bundles three of its own models (Raven perception, Sparrow turn-taking, Phoenix rendering) into one billed conversation minute — the opposite bet: the minute as an indivisible multimodal unit.
Trivia
-
Synthflow (2025-06-24) shows how wide the unbundled minute swings: the same agent can cost roughly $0.11 to $0.24 per minute depending on stack picks — GPT-4.1 mini adds $0.02/min, full GPT-4.1 adds $0.05/min, and bringing your own Twilio drops the telephony meter to $0.00/min. The "price" of a Synthflow minute is really a configuration, not a number.
-
Vapi's headline $0.05/min (verified 2026-06-09) is only its hosting fee — model and telephony costs pass through at cost, so real deployments run roughly $0.13–$0.33/min. The advertised rate covers as little as 15% of what a buyer actually pays per minute.
-
Deepgram (2026-05) prices the BYO discount to the cent: its Voice Agent "Standard – BYO TTS" tier is $0.065/min versus $0.075/min fully managed — an exact $0.010/min line item for one swapped component, the clearest published unit price for unbundling in the corpus.
-
Retell AI (verified 2026-06-09) meters calls to the nearest second with no per-call rounding — but the meter keeps running during silence and hold, because the speech-to-text engine stays active and listening. Unbundling makes even dead air a priced component.
-
The bundled camp's clearest member started unbundling. Bland AI — logged here as the vendor that "repriced the bundle rather than unbundling it" — dropped telephony out of its all-in per-minute rate on 2026-07-21. The rate now covers LLM + STT + TTS only, and the FAQ reads "Telephony is billed separately, on your own carrier or Bland's at pass-through cost." Telephony was the first component Synthflow, Retell and Vapi unbundled too; it is the component that cannot be made to look like software.
-
Synthflow's unbundled price list survived 15 days. It published the full component meter on 2026-06-09 — no platform fee, $0.09/min Voice Engine plus LLM $0.02–$0.05/min plus telephony $0.00–$0.02/min, roughly $0.11–$0.24 all-in — and on 2026-06-24 replaced the entire page with a single Enterprise plan from $30,000/year via Contact Sales. It did not re-bundle; it left self-serve. The unbundled minute may be a better model that is harder to sell without a salesperson to assemble it.
-
LiveKit cut a component 80% and its own advertised price did not move. On 2026-08-11 GPT-5.6 Luna fell from $0.0040/min to $0.0008/min and Terra from $0.0101 to $0.0081, with Sol held at $0.0203 — yet the pricing page's default voice-agent estimate still reads $0.0672/min, because that calculation runs on Gemma 4 31B, which was not among the repriced models. Plan fees untouched at Build $0 / Ship $50 / Scale $500. A headline that ignores an 80% component cut is a parts list, not a price.
-
A component in a voice bill of materials can now cost nothing. LiveKit's 2026-07-29 release added Fish Audio TTS including a FREE SKU alongside two Gemini variants — the first $0-rated component in a voice pass-through catalog in this corpus. When one line of the assembly is zero, the floor of an assembled minute collapses toward the platform take alone, which for LiveKit is a flat ~$0.01/min. Nothing a bundled per-minute rate can answer.
-
Unbundling now runs one level deeper than "which vendor". xAI split its single Speech-to-Speech rate on 2026-08-04 into grok-voice-think-fast-1.0 at $0.05/min and grok-voice-think-fast-2.0 at $0.08/min — a 60% premium for a higher grade of the SAME component — plus $0.004 per message of text input. Sarvam's Dubbing API (2026-08-15) does the workflow version, doubling its per-minute rate on an `editor_flow: true` request flag. Assembling a minute means picking a grade and a workflow, not just a supplier.
-
LiveKit's default price changed because its model catalog changed. The 2026-07-21 release added six per-minute LLM SKUs — Kimi K2.6 $0.0035/min, GPT-5.6 Luna $0.0040, Grok 4.3 $0.0042, Grok 4.5 $0.0070, GPT-5.6 Terra $0.0101, GPT-5.6 Sol $0.0203 — and the pricing-page calculator's default total moved from $0.0735/min (a GPT-5.3 stack) to $0.0672/min (a Gemma 4 31B stack). Plan fees were untouched at Build $0 / Ship $50 / Scale $500. When the advertised number is a calculator output, the vendor no longer has a price — it has a bill of materials.
-
Cartesia proved the bundled camp's repricing risk in one edit: on 2026-07-22 it deleted the Monthly/Yearly (Save 20%) toggle, so the page went from defaulting to $4/$39/$239 yearly- equivalent rates to showing monthly-only Pro $5 / Startup $49 / Scale $299. Nothing about the product changed; annual buyers simply lost roughly 20–25%, delivered as a single opaque number with no component to point at.
For buyers
On an unbundled platform the advertised per-minute rate is a floor, not a price — Vapi's $0.05/min headline covers as little as ~15% of a real deployment's $0.13–$0.33/min. Model an actual stack (which LLM, whose telephony, whose TTS) before comparing vendors, and treat the BYO levers as the negotiation: bringing your own Twilio zeroes Synthflow's telephony meter, and Deepgram prices a swapped TTS at exactly $0.010/min off. Check what the meter counts, too — Retell bills to the nearest second but keeps metering through silence and hold. On bundled platforms (Cartesia, Bland, Tavus) compare the all-in minute directly, but expect repricing risk to arrive as one opaque number: Bland's free-tier minute rose 55% in a single rewrite.
For vendors
Running the unbundled play needs component-level metering (per-second resolution at Retell), a rate card per swappable component, pass-through billing that tracks upstream LLM/TTS/telephony prices at cost, and BYO-key support priced as an explicit per-minute delta — Deepgram's $0.065 BYO-TTS vs $0.075 fully-managed pair is the template. The strategic cost is a thin, visible platform fee under constant deflation pressure: ElevenLabs has cut its Agents pricing repeatedly and Synthflow dropped its platform fee to zero, so margin has to come from volume and orchestration value. Bundling stays viable where the buyer wants one predictable number — Tavus deliberately bills three of its own models as a single conversation minute.
Outlook — what to watch
First logged in June 2026 off the wave-27 voice-AI intake, at a corpus of 207 (now 338). Expect the unbundled camp to grow: component costs keep deflating and pass-through pricing transmits those cuts without repricing, while BYO discounts deepen as buyers consolidate their own Twilio and model contracts. The trend would sharpen if a bundled vendor breaks out component meters or published BYO deltas become the norm; it would weaken if buyers reject bill-of-materials quotes and the flat-minute camp (Cartesia, Bland, Tavus) wins on predictability. Watch whether unbundlers add flat all-in SKUs on top of the meters — that would signal the bundle pulling back ahead.
Bottom line
Five voice-agent platforms — Synthflow, Retell, Vapi, Deepgram, ElevenLabs — now price the minute as a thin platform fee plus at-cost component pass-through, with BYO keys as a published discount lever, while Cartesia, Bland and Tavus keep selling one bundled minute. On unbundled platforms the advertised rate is a floor: price your real stack before comparing.
FAQ
Why does my voice AI agent cost more than the advertised per-minute rate?
Because on unbundled platforms the headline number is only the platform fee. Vapi's $0.05/min covers hosting alone — model and telephony pass through at cost, so real deployments run roughly $0.13–$0.33/min. Synthflow's minute swings from about $0.11 to $0.24 depending on stack picks. Always price a configured stack, not the headline.
What does BYO (bring your own) mean in voice AI pricing?
Plugging your own contract in for a component — your Twilio account for telephony, or your own TTS or LLM key — so that line drops off the platform's meter. Bringing your own Twilio zeroes Synthflow's telephony line, and Deepgram publishes the discount to the cent: $0.065/min with BYO TTS versus $0.075/min fully managed.
Which voice AI platforms unbundle the minute, and which bundle it?
In the corpus, Synthflow, Retell AI, Vapi, Deepgram and ElevenLabs price components separately or pass them through at cost. Cartesia, Bland AI and Tavus sell one flat bundled minute — Tavus deliberately meters its three-model pipeline (perception, turn-taking, rendering) as a single indivisible unit.
Does per-second billing mean I only pay for talk time?
No. Retell meters calls to the nearest second with no per-call rounding, but the meter keeps running during silence and hold because the speech-to-text engine stays active and listening. Unbundling makes even dead air a priced component.