AI Summary
About
Cartesia is a San Francisco-based voice AI startup founded in 2024 by Karan Goel and Albert Gu, the latter best known as co-author of the Mamba state-space model paper. The company’s core bet is that state-space models (SSMs) beat transformer architectures for streaming, low-latency audio generation — a thesis embodied in their flagship Sonic text-to-speech model, which advertises sub-90ms model latency for real-time conversational use cases.
Cartesia raised a $27M seed round led by Index Ventures in March 2024, followed by a $64M Series A in March 2025 also led by Index. Lightspeed Venture Partners, Conviction, and a roster of AI researchers participated. The company is private and pre-revenue-disclosure, with industry estimates putting late-2025 ARR somewhere below $50M — small relative to ElevenLabs but growing fast on the back of the voice-agents wave.
The product surface has three pillars: the Sonic TTS API (per-character or per-credit usage), instant and professional voice cloning, and a packaged Voice Agents product that bundles STT, TTS, and turn-detection into a per-minute conversational unit. Cartesia competes directly with ElevenLabs at the top of the market, PlayHT and Resemble AI in the prosumer tier, and Vapi/Retell on the voice-agents side. Their differentiator is latency plus on-premise availability — two attributes that matter to healthcare, financial-services, and regulated enterprise prospects whose SaaS-only competitors cannot serve.
Pricing summary : How Cartesia’s credit-based freemium stack works
Cartesia runs a credit-based freemium subscription for its core TTS and STT products, with a separate flat per-minute price for Voice Agents drawn from a prepaid dollar balance. As of a 2026-08-28 recapture, all four self-serve tiers are billed at a single monthly rate with no annual-discount option published — Free ($0/mo), Pro ($5/mo), Startup ($49/mo), and Scale ($299/mo) — each bundling a fixed credit allowance (20K → 100K → 1.25M → 8M credits/month), a matching prepaid agent balance, and per-product concurrency limits. Every plan includes unlimited workspace seats and voice slots. Above Scale, the Enterprise tier converts to custom credit and agent volumes with negotiated volume pricing.
The pricing page has four billing dimensions running in parallel: (1) the subscription itself, now a flat monthly price with no billing-period choice, (2) credits spent per character of Sonic-3.6 text-to-speech and per second of Ink-2 transcription, (3) voice-agent minutes at $0.06/min funded by a prepaid dollar balance, and (4) telephony minutes at $0.014/min when you use a Cartesia-provided phone number. A Monthly / Yearly (Save 20%) billing-period toggle has appeared and disappeared three times on this page in the last five weeks: present through mid-2026, removed by 2026-07-22 (advertised rates rose to the current $5/$49/$299), briefly reinstated by 2026-08-26 with a discounted Yearly default (Monthly still showed the same $5/$49/$299), and gone again as of this 2026-08-28 capture — no toggle, no “Save 20%” copy, and no annual-billing control anywhere on the page or in the “Compare plans in detail” grid.
The “credit” is a deliberately abstracted unit, and the pricing page never states its conversion rate — you have to go to the docs. Sonic-3.6 text-to-speech burns approximately 1 credit per character, while Ink-2 speech-to-text burns 3 credits per second of audio on the real-time endpoint; localizing a voice costs 225 credits per added voice accent. Cartesia sunset its Voice Changer API (/voice-changer/bytes, /voice-changer/sse) effective 2026-08-20 — the feature’s former 15-credits-per-second rate no longer applies, and the Pro Voice Clone per-character multiplier and fine-tune cost, last published around 1.5 credits per character and 1,000,000 credits per fine-tune, are no longer shown on Cartesia’s live pricing or docs pages as of this 2026-08-28 revalidation. The pricing page only translates the allowance into time for you: ~27 / ~133 / ~1,667 / ~10,667 TTS minutes per month across Free → Scale, and ~1h 51m / ~9h 16m / ~115h 44m / ~740h 44m of Ink-2 speech-to-text — implying about 12.5 credits per second of speech, which is exactly ~1 credit per character at a normal speaking rate. This is similar to how credit-based billing and other credit-based AI pricing models trade transparency for flexibility: Cartesia can tune per-feature credit rates without renegotiating contract pricing, but customers cannot easily forecast bills without modelling their feature mix.
Voice Agents (the Line product, labeled “Managed Agents” in the pricing-page comparison grid) are billed separately on media-minute pricing rather than in credits: a flat $0.06 per minute of call duration on every plan, Free through Scale. That is a deliberate packaging choice that simplifies forecasting for conversational use cases at the cost of obscuring which component (Ink-2 STT, Sonic-3.6 TTS, or turn-detection) drives the unit cost. This dual-axis model — credits for batch TTS, dollars-per-minute for real-time agents — mirrors the broader shift from per-user licensing to usage-based AI billing playing out across the category.
What makes this different: Cartesia is one of the few voice AI vendors advertising on-premise Sonic and Ink deployment in its Enterprise tier. ElevenLabs, OpenAI Voice, and PlayHT are all SaaS-only. That single capability — the models running inside a hospital’s or bank’s network — is the wedge that justifies enterprise pricing above the self-serve ceiling.
Pricing by product
Sonic-3.6 TTS / Ink-2 STT (self-serve credit tiers)
| Tier | Price | Included | Key mechanics |
|---|---|---|---|
| Free | $0/mo | 20K credits/mo (~27 TTS min, ~1h 51m STT); $1/mo prepaid agents; 1 agent slot | No credit card; unlimited workspace seats and voice slots; Sonic-3.6 + Ink-2 |
| Pro | $5/mo | 100K credits/mo (~133 TTS min, ~9h 16m STT); $5/mo prepaid agents; 3 agent slots | ”Everything in Free, plus” commercial use license and instant voice cloning |
| Startup | $49/mo | 1.25M credits/mo (~1,667 TTS min, ~115h 44m STT); $49/mo prepaid agents; 5 agent slots | ”Everything in Pro, plus” pro voice cloning and Organizations |
| Scale | $299/mo | 8M credits/mo (~10,667 TTS min, ~740h 44m STT); $299/mo prepaid agents; 10 agent slots | Self-serve ceiling; “Everything in Startup, plus” priority support and high concurrency limits |
Concurrency is metered separately per product: 2 / 3 / 5 / 15 concurrent TTS requests, 8 / 12 / 20 / 60 concurrent STT requests, and 8 / 12 / 20 / 60 concurrent voice-agent calls across Free → Scale. As of a 2026-08-28 recapture, the pricing page has no Monthly / Yearly toggle at all — every tier shows a single flat monthly rate ($5 Pro, $49 Startup, $299 Scale) and the “Compare plans in detail” grid subtitle no longer carries its former “same billing period as above” qualifier. This is a reversal of the 2026-08-26 state, when the toggle was present and defaulted to a discounted Yearly view; it had also been absent between 2026-07-22 and roughly 2026-08-26. Credit allowances, prepaid agent balances, and concurrency limits have been unchanged across all three states.
Enterprise (custom)
| Tier | Price | Included | Key mechanics |
|---|---|---|---|
| Enterprise | Custom (Contact us) | Custom credits & agent usage; custom concurrency limits | Volume pricing; “Everything in Scale, plus” DPAs and BAAs for compliance, shared Slack channel, SSO, security questionnaires |
Voice Agents — Line (flat per-minute, all plans)
| Component | Rate | Notes |
|---|---|---|
| Call duration | $0.06 per minute | Identical flat rate on Free, Pro, Startup, and Scale; Custom on Enterprise |
| Telephony | $0.014 per minute | Only “when using a Cartesia-provided phone number” |
| Agent funding | Prepaid dollar balance | $1/mo (Free) → $5 (Pro) → $49 (Startup) → $299 (Scale) |
| LLM usage during calls | Included | ”Free for UI-created agents for a limited time” — the caveat is on the page, so budget for it to end |
| Evaluations | Included | ”Free of charge for a limited time only” |
The “Compare plans in detail” grid on the pricing page relabeled its Voice Agents section header from “Line” to “Managed Agents” and renamed the “Number of agent slots” row to “Cartesia-provisioned phone numbers” between the 2026-08-26 and 2026-08-28 captures; the underlying product is still called Line on docs.cartesia.ai, and the row’s values (1 / 3 / 5 / 10 / Custom across Free → Scale → Enterprise) are unchanged.
Voice cloning & feature credits
| Service | Cost | Notes |
|---|---|---|
| Instant voice cloning | Included | From the Pro tier up ($5/mo); not available on Free |
| Professional voice cloning | Included | Introduced at the Startup tier ($49/mo); not available on Free or Pro |
| Voice changer | Discontinued | Cartesia sunset the /voice-changer/bytes and /voice-changer/sse endpoints effective 2026-08-20 (per Cartesia’s own API reference); the former 15-credits-per-second rate no longer applies and the feature is absent from both the pricing page and the docs pricing summary as of this 2026-08-28 revalidation |
| Localizing a voice | 225 credits per added voice accent | Per-voice localization, same rate on every tier; page copy changed from “as a one-time cost” to “per added voice accent” between the 2026-08-26 and 2026-08-28 captures, credit amount unchanged |
| Pro Voice Clone fine-tune | ~1,000,000 credits per fine-tune (unverifiable as of this revalidation) | Last published in the docs through at least 2026-08-14; no longer listed on the docs pricing summary or the pricing page as of 2026-08-26 onward — treat as a historical figure, not a currently confirmed one |
Credit consumption rates (published in the docs, not on the pricing page)
| Operation | Credit rate |
|---|---|
| Sonic-3.6 text to speech | ~1 credit per character |
| Sonic-3.6 text to speech, Pro Voice Clone voice | ~1.5 credits per character (last published; no longer listed on docs.cartesia.ai/pricing as of 2026-08-26 onward — unverifiable at this revalidation) |
| Ink-2 speech to text (real-time) | 3 credits per second of audio |
| ink-whisper speech to text (real-time) | 1 credit per second of audio |
| ink-whisper speech to text (batch) | 1 credit per 2 seconds of audio |
| Infill | 300 credits + ~1 credit per character |
Voice changer (formerly 15 credits per second of input audio) has been removed from this table: Cartesia sunset the /voice-changer/bytes and /voice-changer/sse endpoints effective 2026-08-20, per its own API reference documentation, and no credit rate for the feature appears on any current Cartesia surface.
The pricing page publishes only the time-translated allowance (“~133 minutes”), never the credit rate behind it — so the ~1-credit-per-character conversion above is the number a buyer actually needs to model a bill, and it lives one host away on docs.cartesia.ai/pricing.
Sales motions across products: self-serve PLG for Free, Pro, Startup, and Scale (all four are dashboard-purchasable, including Voice Agents); sales-led only for the custom Enterprise tier.
Hidden costs : What Cartesia users actually pay beyond the headline credit allowance
Archetype A: Solo developer building a voice-cloned podcast workflow (Startup plan)
A solo creator generating ~5 hours of narration per month using instant voice cloning:
| Line item | Monthly cost |
|---|---|
| Startup subscription ($49/mo) | $49.00 |
| Base allowance: 1.25M credits (~1,667 TTS min) | included |
| ~5 hrs (300 min) of TTS — well within ~1,667-min allowance | within allowance |
| Localizing one voice into a second language (225 credits, one-time) | within allowance |
| Estimated total | ~$49 |
For a pure-TTS creator, the Startup plan’s ~1,667 included minutes are hard to exhaust at podcast volumes, so the headline $49 is usually the whole bill. The forecasting trap is feature credits, not volume: the voice changer alone burns 15 credits per second of audio, so heavy use of premium features — not raw narration length — is what pushes a user toward overage. This is the most common form of AI cost unpredictability on credit-based platforms.
Archetype B: Mid-market customer-service team running voice agents (Scale plan)
A 50-agent contact-center deployment, average 4 hours/day of conversational voice across the team:
| Line item | Monthly cost |
|---|---|
| Scale subscription ($299/mo) | $299 |
| Voice-agent minutes: 50 agents × 4 hrs × 20 days = 240,000 min × $0.06 | $14,400 |
| Telephony (Cartesia numbers): 240,000 min × $0.014 | $3,360 |
| TTS credits for IVR + scripted prompts (within 8M Scale allowance) | included |
| Estimated total (approximately) | ~$18,000 |
At 240k minutes, the $299 base subscription is rounding error. Real cost sits in the flat $0.06/min agent rate plus $0.014/min telephony, both drawn from the prepaid agent balance. Because the per-minute rate is identical on every plan, the only lever at this volume is the custom Enterprise tier, where Cartesia advertises volume pricing on credits and agent usage.
Want to estimate your own Cartesia bill? Use the Cartesia pricing calculator to model costs across voice cloning, TTS character volume, and agent-minute consumption.
Pricing evolution : Cartesia’s journey from pure-usage API to credit-based SaaS
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2024 Q1 | 0 | 1 | Company launch + Sonic announced; $27M seed |
| 2024 Q2 | 0 | 1 | Sonic API GA; pure usage at ~$65/1M chars |
| 2024 Q3 | 1 | 2 | Tiered subscriptions introduced: Free, Creator ($5), Pro ($49) |
| 2025 Q1 | 1 | 3 | Sonic-2 launch; clone multiplier formalized; $299 production tier added; Voice Agents beta announced alongside the Series A |
| 2025 Q3 | 0 | 2 | Scale + Enterprise tiers formalized; on-prem SKU added |
| 2026 Q1 | 1 | 1 | Voice Agents GA; flat $0.06/min call duration published; prepaid agent dollar balances replace bundled minutes |
| 2026 Q3 | 3 | 0 | Monthly / Yearly (Save 20%) toggle flips three times — removed 07-22, reinstated 08-26, removed again 08-28 — plus Voice Agents relabeled “Line” → “Managed Agents” |
Tracked range: 2024 Q1–2026 Q3. Quarters not listed verified stable (0 changes, 0 additions). Sonic-2 voice-clone multiplier (2×) is technically a price change disguised as a product launch.
Notable changes
- 2024-09-30 — Moved from pure-usage API ($65/1M chars) to a freemium subscription stack. First major pricing-architecture decision, made roughly six months post-launch.
- 2025-01-22 — Sonic-2 launch silently introduced the 2× credit multiplier for cloned voices. No public announcement; documented in API reference only. This is the kind of packaging change that benefits the vendor and surprises customers.
- 2025-03-25 — Voice Agents launched as a separate per-minute SKU rather than rolled into the credit system — a packaging decision that has held since.
- 2025-09-15 — Scale tier and on-prem Enterprise SKU formalized. This was the first explicit sales-led motion at Cartesia and signaled enterprise readiness.
- 2026-02-10 — Voice Agents GA at a flat $0.06 per minute of call duration on every tier, funded from a prepaid dollar balance rather than a bundle of included minutes. Publishing one rate for Free through Scale removed the usual per-plan rate card — and with it any volume incentive to climb the ladder for agent work.
- 2026-07-22 — The “Self-serve plans & usage” Monthly / Yearly (Save 20%) toggle disappeared from the pricing page. Nothing else moved: allowances, prepaid agent balances, concurrency, and the $0.06/min agent rate all held. The advertised entry prices rose to $5 / $49 / $299 because the discounted yearly view — $4 / $39 / $239 — was the default the page had been showing.
- 2026-08-26 — The toggle removed five weeks earlier reappeared, again defaulting to Yearly ($4/$39/$239 vs. $5/$49/$299 Monthly). The flagship model was also relabeled Sonic-3.6 on the pricing and product pages (from Sonic-3.5) — a version bump, not a pricing change.
- 2026-08-28 — The toggle disappeared again, two days after returning — the third state change to this single control in five weeks. Advertised rates reverted to the flat $5/$49/$299 monthly-only figures. Two copy relabels landed alongside it: Voice Agents’ “Line” section became “Managed Agents” (and “Number of agent slots” became “Cartesia-provisioned phone numbers”), and voice localization’s “one-time cost” framing became “per added voice accent” — both cosmetic, with values and credit amounts unchanged.
The yearly-billing toggle’s three reversals in detail
What started as a single removal has become a pattern: the Monthly / Yearly (Save 20%) toggle has now flipped three times in five weeks — present through mid-2026, removed 2026-07-22, reinstated 2026-08-26 (defaulting to Yearly again), and removed a second time 2026-08-28. A prospect who checked the pricing page in late June, again in late July, again on 26 August, and once more two days later would have seen the advertised Pro price move $4 → $5 → $4 → $5.
Three things make this more than a cosmetic edit, and the repeat sharpens each of them:
- There is still no committed-term price anywhere below Enterprise, and the control keeps arriving and leaving without warning. Annual prepay is the only mechanism a self-serve buyer has ever had to trade flexibility for cost on this page, and it has now proven unreliable enough — on, off, on, off — that neither state can be treated as “the” price. A budget built off a screenshot risks being wrong within 48 hours.
- The discount never applied to the whole bill in either appearance. Both before mid-2026 and during the 2026-08-26 reinstatement, the Yearly view discounted the subscription line only ($4 vs. $5 for Pro) while the prepaid agent balance, credits, and the $0.06/min call rate stayed flat. Toggling it on or off therefore does little for heavy agent users and moves the number that matters most for pure-TTS subscribers.
- All three flips shipped silently. No announcement, no changelog entry, and no grandfathering language accompanied the July removal, the August reinstatement, or the August re-removal. The only visible trace each time is a vestigial “same billing period as above” string under the comparison grid that appears when the toggle exists and vanishes when it doesn’t — a page-state artifact, not a communicated pricing decision.
The read has shifted since July. A single quiet removal twelve months after a funding round could plausibly pass as confidence — a company that doesn’t need annual prepay for cash flow deleting a discount it doesn’t need. Three flips in five weeks reads differently: either Cartesia is actively testing whether the annual discount moves conversion (an experiment run directly on the live pricing page, with real prospects as the test group), or pricing-page edits are shipping without the coordination a public price should have. Either way, the pattern itself — not the direction of any single flip — is now the more important signal for a buyer deciding whether to trust what they see on the page today.
What’s unique : Cartesia’s distinctive pricing mechanics
1. Credits as a per-model multiplier abstraction. Cartesia’s credit unit lets the company adjust per-model costs (approximately 1.5 credits per character for Pro Voice Clone voices against approximately 1 for standard Sonic-3.5; 3 credits per second for Ink-2 speech-to-text against 1 for ink-whisper) without changing the headline subscription price. This protects margin on premium models but creates aggregation complexity for customers who mix voices. ElevenLabs uses a simpler character-only model; PlayHT uses words. Cartesia’s credit abstraction is closer to OpenAI’s old “tokens” model in spirit — flexible for the vendor, harder to forecast for the buyer.
2. Dual-axis billing: credits for TTS, minutes for Voice Agents. Most voice AI vendors bill everything in the same unit (characters or minutes). Cartesia uses credits for batch TTS and minutes for agents — recognizing that a 2-minute customer-service conversation is operationally different from a 500-character notification. This composite billing approach aligns price to use case but doubles the forecasting work for buyers who do both.
3. On-premise as the enterprise wedge. Cartesia is the only major voice AI vendor offering on-prem Sonic deployment for healthcare and financial-services prospects. ElevenLabs is SaaS-only. OpenAI Voice is SaaS-only. PlayHT is SaaS-only. This single capability supports a separate Enterprise SKU with custom pricing — converting a regulatory constraint into a revenue line. See Deepgram’s blueprint for a comparable speech-side play.
4. Instant Voice Clone bundled into the $5 prosumer tier, with unlimited voice slots. Instant voice cloning starts at Cartesia’s $5 Pro tier — effectively level with ElevenLabs, whose $6/mo Starter plan is also where instant cloning begins — but Cartesia advertises unlimited workspace seats and voice slots on every plan rather than metering how many clones you may keep. Professional cloning is where the two diverge in the other direction: Cartesia gates it to the $49 Startup tier, while ElevenLabs opens it at $22/mo Creator. The prosumer play is therefore about slot generosity and a $5 entry price, betting that creators who clone voices at $5 upgrade to Startup ($49) as they scale — a textbook PLG monetization funnel, though the entry-price advantage over ElevenLabs has effectively closed.
5. The annual-discount toggle has become the least stable part of the page. The Monthly / Yearly (Save 20%) control has flipped three times in five weeks — removed 2026-07-22, reinstated 2026-08-26 (again defaulting to Yearly), removed a second time 2026-08-28 — while every other number on the page (credit allowances, agent balances, the $0.06/min call rate) held constant across all three states. Most infrastructure vendors run the opposite play: pick a discount and hold it to pull buyers onto annual terms. Cartesia’s version of that lever now looks less like a settled “month-to-month by choice” position and more like an unresolved question about whether the discount earns its keep — plausible given the $64M Series A that removed any pressure to buy cash flow with prepay. Whichever it is, the trade for a buyer is the same each time the toggle is off: list prices sit ~25% higher (Scale’s effective rate reaches ~$37 per 1M credits) with no committed-term alternative published.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Generous free tier (20K credits, ~27 TTS min) builds developer goodwill without credit card friction | Credit multipliers and per-feature rates (voice changer at 15 credits/sec) are documented but easy to miss — predictable cause of bill shock |
| Sub-90ms latency is genuinely category-leading and supports real-time use cases | The yearly-billing toggle has flipped three times in five weeks (removed 07-22, back 08-26, removed again 08-28) — whether a committed-term discount exists now depends on which day you load the page |
| On-prem Enterprise SKU is rare in voice AI and supports HIPAA/SOC 2 use cases competitors cannot serve | None of the three toggle flips (07-22, 08-26, 08-28) shipped with an announcement or grandfathering language — the only trace each time is a vestigial “same billing period as above” line that appears and disappears with the control |
| Dual-axis pricing (credits + minutes) maps to actual use cases (batch vs conversational) | A flat $0.06/min agent rate on every tier means high-volume agent buyers get no volume relief until custom Enterprise pricing |
| Whole ladder to $299/mo is self-serve and month-to-month — “Pause or cancel anytime”, no term commitment to reach the ceiling | No published education or nonprofit discount tier; missing a key prosumer wedge ElevenLabs uses well |
| Startup ($49) tier delivers ~1.25M credits (~1,667 TTS min) — roughly 10× the monthly allowance of ElevenLabs Creator ($22/mo for 121k credits) for a little over 2× the price | Credit-unit obfuscation makes apples-to-apples comparisons against character-priced competitors hard for buyers |
Billing UX : Cartesia’s account controls and developer console
- Self-serve plan selection — Every tier below Enterprise has its own in-page action: “Start free”, “Select Pro”, “Select Startup”, “Select Scale”. Only Enterprise routes to “Contact us”, so the whole ladder up to $299/mo is dashboard-purchasable without a sales call.
- “Calculate your costs based on your usage needs” — An on-page cost calculator with a two-step flow: (1) select features and capabilities (Text to Speech, Speech to Text, Voice Agents) and (2) drag a “Minutes of generated audio per month” slider from 0 to 50,000+. It returns a “Recommended plan” card with the price and the included TTS minutes (e.g. 45 min/month → Pro, $5/mo, 133 Text to Speech minutes/month).
- “Compare plans in detail” — A full feature-by-feature matrix across Free, Pro, Startup, Scale, and Enterprise, broken out by product (Text to Speech / Sonic-3.6, Speech to Text / Ink-2, Voice Agents / Managed Agents [Line]) with included minutes, concurrency, and per-feature credit rates in one grid. As of the 2026-08-28 recapture its subtitle reads just “Feature-by-feature across Free, Pro, Startup, and Scale.” — the earlier ”— same billing period as above” qualifier is gone, matching the toggle’s removal.
- No billing-period toggle (currently) — The pill-style Monthly / Yearly (Save 20%) toggle that sat above the plan cards is absent as of this 2026-08-28 recapture; every tier now shows one flat monthly price (Pro $5, Startup $49, Scale $299). This control has been on-again/off-again: present through mid-2026, removed by 2026-07-22, back with a discounted Yearly default as of 2026-08-26, and gone again by 2026-08-28.
- “Pause or cancel anytime” — Stated directly on the recommended-plan card, alongside published FAQ entries for “What if I cancel or downgrade my tier in the middle of a payment period?” and “When does my subscription renew?”.
- Rollover credits and tier changes — Documented FAQ mechanics for “What happens to my rollover credits if I change my pricing tier?” and “What happens if I upgrade to a higher subscription tier?” — so credit balances, not just prices, move when you change plan.
- Overage handling on two separate balances — The FAQ splits overage into “What happens if I use more model credits than I have in my account?” and “What happens if I use more prepaid voice agent dollars than I have in my account?”, plus “How and when are overages charged?” — confirming credits and agent dollars drain independently.
- “Toggle overages” on the subscription page — Newly documented on
docs.cartesia.ai/pricingas of the 2026-08-28 recapture: with overages enabled, requests keep succeeding past the monthly credit allotment and the excess bills as an overage (the credit balance can go negative); with overages disabled, requests that would exceed the allotment fail until renewal or upgrade. The control lives on the account’s subscription page. - Break-tag billing rule — A published FAQ entry, “How are break tags counted in billing?”, covers how SSML break tags are metered against credits, a common source of surprise on per-second TTS.
- Compliance and deployment controls — Enterprise adds DPAs and BAAs, SSO, security questionnaires, and a shared Slack channel; the page’s FAQ separately answers “Is Cartesia SOC 2 Type II certified?” and “Can I deploy the models on-premise or in a virtual private cloud?” — useful when aggregating consumption across teams under one contract.
Strategic wins : Why Cartesia’s pricing decisions worked
1. Free tier with no credit card removed friction at the top of the funnel
20,000 free credits monthly — roughly 27 minutes of Sonic-3.5 audio, enough to actually evaluate the model in a real prototype — combined with no credit card requirement created one of the most frictionless voice AI onboarding flows in the category. This is the same playbook ElevenLabs used in 2023 (and has since tightened); Cartesia is running it harder. The PLG-style free-tier play is the dominant reason Cartesia shows up on indie-developer benchmark posts more often than its market share would suggest.
2. $5 Pro tier captured prosumer segment competitors ignored
Pricing Pro at $5/month, below the typical prosumer floor of $20, opened a segment ElevenLabs and PlayHT had effectively abandoned. The bet was that $5/month creators eventually upgrade to $49 Startup as their projects grow — a classic tiered monetization ladder. Even if the conversion rate is low, the LTV math works because $5 tier is profitable at moderate volumes given the credit allowance.
3. On-prem Enterprise SKU unlocked regulated-industry revenue
Cartesia’s decision to offer on-premise Sonic deployment for HIPAA-regulated healthcare prospects converted a technical capability (model portability via SSM efficiency) into a revenue line. ElevenLabs, OpenAI Voice, and PlayHT cannot serve these customers at all. This is a textbook example of turning compliance constraints into pricing power — the contracts are large, sticky, and effectively uncontested.
4. Series A timing locked in pricing before commoditization pressure
Raising $64M just 12 months after seed gave Cartesia runway to hold pricing while transformer-based TTS commoditized. Most competitors are now forced into per-character price wars; Cartesia can hold the line on credit pricing and let cheaper alternatives compete for the prosumer floor. The clearest evidence is the toggle’s own trajectory: withdrawn 2026-07-22 (a ~25% effective list increase with no new capability attached), reinstated 2026-08-26, then withdrawn again 2026-08-28. A vendor defending margin from a weak cash position discounts to pull revenue forward and holds that line; a vendor with runway can afford to toggle the lever on and off testing conversion without the reversal threatening anything it actually needs. The funding-pricing relationship is rarely discussed but is a plausible reason Cartesia can treat its own list price as adjustable inventory rather than a fixed commitment.
Areas to improve : Gaps in Cartesia’s pricing approach
1. Credit multiplier opacity creates predictable bill shock
The ~1.5× credit cost for Pro Voice Clone voices — approximately 1.5 credits per character against approximately 1 for standard Sonic-3.5 — is published on docs.cartesia.ai/pricing but never appears on the pricing page itself. A user planning around the headline “8M credits” on Scale who switches to Pro Voice Clone voices mid-month will hit overage at roughly 5.3M characters, not 8M. The fix: surface the effective character-equivalent allowance for each plan and add a real-time projection in the usage dashboard. See bill shock patterns in AI billing for why this matters more than it looks.
2. A toggle that appears and disappears is worse than no toggle at all
The Monthly / Yearly control has now been removed (2026-07-22), reinstated (2026-08-26), and removed again (2026-08-28) — three states in five weeks, none of them announced. Restoring the toggle briefly proved Cartesia can ship the committed-term option; removing it again two days later proved it won’t stay. Finance teams that want a fixed annual number for a $299/mo workload can no longer trust a screenshot of the pricing page, let alone plan around it, and still have to open a sales conversation about a tier they do not otherwise need. The fix is not another toggle flip: publish an annual-prepay price on Startup and Scale as an explicit, versioned line (“$3,588/yr, or $2,990 prepaid”) that survives independently of whichever way the toggle happens to be pointed this week, and pair any future change to it with a changelog entry. A discount that disappears without explanation costs trust once; one that disappears three times costs it on a recurring basis.
3. No published rung between $299 and “contact us”
The jump from Scale ($299/mo, 8M credits, ~10,667 TTS minutes) to custom Enterprise is a 1-step cliff, and the July 2026 change removed the one thing that used to soften it. A team consuming 30-100M credits per month has no published price and no prepay option — only a sales call whose outcome they cannot benchmark. The same cliff exists on the agent side, where $0.06/min is flat from Free through Scale with no volume break until Enterprise. The fix: introduce a published $999-$1,500 tier with a stated agent-minute rate — a missing rung in the pricing ladder that competitors will otherwise fill.
4. Feature credit costs live in a collapsed FAQ, not on the plan cards
The pricing page publishes exactly two per-feature credit rates in the comparison grid — the voice changer at 15 credits per second of audio, and voice localization at 225 credits one-time. Everything else a buyer needs to size an allowance sits behind accordion questions: “How many credits do I need?”, “How many credits do I need for Pro Voice Cloning?”, “How are break tags counted in billing?”. Professional voice cloning is included from the $49 Startup tier, but its credit draw — the number that decides whether 1.25M credits is generous or tight — is never shown next to the tier that includes it. The fix: put a per-feature credit table on the plan cards themselves, so the allowance and the things that consume it appear in the same view.
5. No education or nonprofit pricing
Most AI infrastructure vendors offer 50% education or nonprofit discounts as a brand-equity and acquisition play. Cartesia has neither published. The fix is mechanical: a SheerID-verified education tier at $25/month (50% off Startup) would capture student researchers and university labs at near-zero cost to Cartesia.
Monetization stack & signals : how Cartesia builds & buys its revenue engine
Buys 2 Builds 0 9 open roles
Cartesia buys its CRM — Salesforce plus HubSpot — and its first GTM-engineering hire builds the pipelines between them. Hiring skews heavily to customer-success and forward-deployed roles, signaling a sales-led enterprise motion grafted onto a self-serve PLG core.
-
“Build reliable data pipelines and integrations across systems including Salesforce, HubSpot, and internal platforms”
-
“Build reliable data pipelines and integrations across systems including Salesforce, HubSpot, and internal platforms”
-
“Experience evaluating or implementing ERP systems at a high-growth company (NetSuite, Sage Intacct, or similar)”
- Forward Deployed Engineer Customer success seen Jun 9, 2026
- Founding Forward Deployed Engineer (India) Customer success seen Jun 9, 2026
- Technical Account Manager Customer success seen Jun 9, 2026
- Technical Account Manager, India Customer success seen Jun 9, 2026
- Strategic Partnerships Manager Customer success seen Jun 9, 2026
- GTM Engineer RevOps seen May 13, 2026
- International Growth Lead Growth seen May 13, 2026
- Solutions Engineer, Pre-Sales Growth seen May 13, 2026
- Account Executive Growth seen May 13, 2026
Signals reviewed · derived from public job posts
Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.
Key takeaways
-
Credit-based billing trades transparency for vendor flexibility — and customers notice. Cartesia’s per-model credit multipliers (~1.5 credits per character for Pro Voice Clone voices, 3 credits per second for Ink-2 speech-to-text) let the company tune margins without contract renegotiation, but they create forecasting friction for buyers. Pricing teams should weigh the operational flexibility of opaque units against the trust cost paid when customers hit unexpected overages.
-
Dual-axis billing (credits + minutes) correctly recognizes that batch and real-time AI workloads are economically different. Cartesia bills TTS by credits and Voice Agents by minutes because the cost drivers differ — GPU inference time for streaming agents dominates over raw token output. As AI products incorporate more real-time conversational surfaces, this dual-axis pattern will become standard.
-
On-premise deployment is the enterprise moat in voice AI. Cartesia’s on-prem Enterprise SKU serves customers that SaaS-only competitors literally cannot. For any AI infrastructure vendor whose model can run efficiently on customer hardware, a premium on-prem tier is one of the highest-margin revenue lines available — and a defensible position against commoditization.
-
Free tiers without credit cards drive disproportionate developer mindshare. Cartesia’s no-CC, 20K-credit free tier is a small giveaway with outsized funnel impact: indie developers benchmark Cartesia first because it costs nothing to try. The conversion cost per qualified evaluator is roughly two orders of magnitude lower than paid marketing for AI infrastructure.
-
A discount that flip-flops erodes more trust than one that simply disappears. Cartesia’s yearly-billing toggle has been removed (2026-07-22), reinstated (2026-08-26), and removed again (2026-08-28) — three unannounced state changes to the same control inside five weeks, each moving the advertised Pro price between $4 and $5. A single quiet removal reads as repricing; three reversals with no changelog read as instability, and buyers price both interpretations as a reason to distrust the page. If you’re testing a billing-period discount, say so — an explicit “currently testing annual pricing” note costs nothing and converts a credibility problem into a transparent experiment.
UBP implications
-
Composite billing units (credits, minutes) abstract pricing risk to the vendor. Cartesia’s credit unit lets the company shift per-model economics without contract changes — a powerful UBP design when the underlying cost structure evolves (new models, new hardware). Pricing teams building usage-based products should evaluate whether a composite unit serves their margin protection better than a transparent single-dimension unit.
-
Real-time conversational AI requires per-minute (not per-token) pricing to align with cost. Voice Agents bill by minute because GPU inference time, network round-trips, and turn-detection compute dominate the cost stack — not raw output tokens. As AI products add real-time surfaces (voice, video, interactive agents), pricing teams should expect to move from per-token to per-second/per-minute billing for these workloads.
-
On-prem and VPC deployment is a defensible UBP wedge for regulated industries. Cartesia’s Enterprise tier proves that customers in healthcare, financial services, and government will pay a multi-X premium for on-prem AI that meets data-residency requirements. For any AI infrastructure provider whose model is portable, an on-prem SKU at 2-5× SaaS pricing represents one of the cleanest enterprise UBP plays available.
Sources
- Cartesia pricing page (accessed 2026-08-28)
- Cartesia documentation home (accessed 2026-08-28)
- Cartesia Sonic model overview (accessed 2026-08-28)
- Cartesia blog — Sonic-2 launch (accessed 2026-05-29)
- Cartesia docs — Pricing (how Cartesia charges for usage; per-character and per-second credit rates) (accessed 2026-08-28)
- Cartesia docs — full documentation index (confirms the Voice Changer API sunset date of 2026-08-20) (accessed 2026-08-28)
- Cartesia docs — Concurrency and scaling for Line voice agents (accessed 2026-07-22)
- Cartesia docs — Text-to-speech API reference (current Sonic model lineup: sonic-3.5, sonic-3, sonic-latest) (accessed 2026-07-22)
- Comparable: ElevenLabs blueprint on UsagePricing (accessed 2026-05-29)
- Comparable: Deepgram blueprint on UsagePricing (accessed 2026-05-29)
Bottom line
Cartesia has built a credit-based freemium stack that competes effectively with ElevenLabs on price at the prosumer tier and with Deepgram on capability at the regulated-enterprise tier. The pricing architecture is clever — credits abstract model differences, dual-axis billing aligns to use case, and on-prem deployment unlocks regulated industries that SaaS-only competitors cannot serve. But the same credit abstraction that protects margin creates forecasting friction; the missing tier between $299 and “call sales” leaves a gap competitors will fill; and the 20%-off yearly toggle has now been removed, reinstated, and removed again inside five weeks (2026-07-22, 2026-08-26, 2026-08-28) — three unannounced page edits that turned the discounted rate into a moving target rather than a stable committed-term option. The bet on state-space architecture for latency-critical voice is structurally sound, and a fast-follow Series A gives Cartesia room to hold positioning while transformers commoditize underneath them. The open question is no longer whether the pricing works, but whether a company still deciding how to price its own annual discount will stay easy to plan around.
Browse the full pricing blueprint to compare Cartesia against ElevenLabs, Deepgram, and other voice AI vendors.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
Yearly Billing Toggle Removed Again, Two Days After Returning
The Monthly / Yearly (Save 20%) toggle reinstated on 2026-08-26 is gone again — the third state change to this single control in five weeks. The pricing page reverts to a single flat monthly rate per tier (Pro $5, Startup $49, Scale $299) with no annual-discount control anywhere on the page. Two copy changes landed alongside it: the "Compare plans in detail" grid relabeled its Voice Agents section from "Line" to "Managed Agents" and its agent-slot row to "Cartesia-provisioned phone numbers" (values unchanged), and voice localization is now described as "225 credits per added voice accent" rather than "as a one-time cost" (credit amount unchanged).
Yearly Billing Toggle Reinstated
The Monthly / Yearly (Save 20%) toggle removed from cartesia.ai/pricing on 2026-07-22 is back. The page now defaults to the Yearly view again — Pro $4/mo, Startup $39/mo, Scale $239/mo — with a Monthly tab showing $5/$49/$299. Credit allowances, prepaid agent balances, concurrency limits, and the $0.06/min voice-agent rate are unchanged. The flagship TTS model is now labeled Sonic-3.6 on the pricing page (was Sonic-3.5).
Yearly Billing Option Removed from the Pricing Page
Cartesia dropped the "Self-serve plans & usage" Monthly / Yearly toggle (Yearly was marked "Save 20%") from cartesia.ai/pricing. The page now advertises a single monthly rate per tier — Pro $5/mo, Startup $49/mo, Scale $299/mo — where the previous default view showed the discounted yearly-equivalent rates of $4, $39, and $239. Credit allowances (20K / 100K / 1.25M / 8M), prepaid agent balances, concurrency limits, and the $0.06/min voice-agent rate are all unchanged.
Voice Agents GA + Flat Per-Minute Pricing
Voice Agents (Line) moved to general availability with a flat $0.06 per minute of call duration on every plan, plus $0.014 per minute of telephony when using a Cartesia-provided phone number. Agent usage is funded from a prepaid dollar balance ($1/mo on Free up to $299/mo on Scale) rather than a bundle of included minutes.
Scale Tier + Enterprise Deployment Introduced
An 8M-credit Scale tier was positioned as the self-serve ceiling with priority support and high concurrency limits, while a custom Enterprise SKU added on-premise / VPC deployment of the Sonic and Ink models with DPAs, BAAs, and SSO for compliance-driven prospects.
Series A ($64M) + Voice Agents Beta
Index Ventures led a $64M Series A. Cartesia simultaneously announced Voice Agents — a packaged STT+TTS+turn-detection product priced per minute rather than per character — positioning against Vapi and Retell.
Next-generation Sonic + Pricing Refresh
Cartesia shipped a next-generation Sonic release with improved naturalness and broad multilingual support, and rebalanced credit consumption so that premium features (voice changer, voice localization) draw additional credits beyond base synthesis. An 8M-credit production tier was added for higher-volume teams.
Tiered Subscription Plans Introduced
Cartesia restructured from pure usage to a freemium subscription model with three tiers: Free (10k credits), Creator ($5/mo for 100k credits), and Pro ($49/mo for 1.25M credits). Credit-pack overages preserved for power users. This was the first move toward a SaaS-style commitment model.
Sonic API General Availability
Sonic TTS API opened to the public with a free tier of 10,000 monthly credits and a developer-focused, pay-as-you-go usage model. No subscription tiers yet — purely metered API with prepaid credit packs. Targeted at developers integrating real-time voice into apps.
Company Launch + $27M Seed Round
Cartesia emerged from stealth with a $27M seed round led by Index Ventures. Founders Karan Goel and Albert Gu announced Sonic, the first commercial state-space TTS model. Initial product was API-only with usage-based per-character pricing at approximately $65 per 1M characters — undercutting ElevenLabs by roughly 30%.
- · Cartesia was founded in 2024 by Karan Goel and Albert Gu — the same Albert Gu who co-authored the Mamba state-space model paper at CMU. Cartesia's Sonic model is a direct commercial application of state-space architecture, betting that SSMs beat transformers for real-time streaming audio.
- · Sonic was the first commercial TTS model to advertise sub-90ms model latency — roughly 3-5× faster than ElevenLabs Turbo at launch. That latency number is itself a marketing artifact: it measures only the model, not the network round-trip a developer actually pays for.
- · Cartesia raised a $27M seed in March 2024 led by Index Ventures, then a $64M Series A in March 2025 also led by Index — an unusually fast follow-on that locked in pricing power before competitors could undercut. Lightspeed, Conviction, and a roster of AI researchers participated.
Questions & answers
- How much does Cartesia cost per month, and is annual billing available?
- Cartesia's self-serve plans are Free (20,000 credits/month), Pro at $5/month (100,000 credits), Startup at $49/month (1.25 million credits), and Scale at $299/month (8 million credits). As of a 28 August 2026 recapture there is no annual-billing option published on the pricing page — no Monthly / Yearly toggle and no "Save 20%" copy anywhere on the page. That control has been on-again/off-again: present through mid-2026, removed by 22 July 2026, reinstated (defaulting to a discounted Yearly view) by 26 August 2026, and gone again two days later. Scale is fully self-serve rather than sales-gated, and above it Enterprise uses custom credit and agent volumes with volume pricing negotiated with sales.
- What is a Cartesia credit and how does it convert to audio?
- Credits are consumed per character of generated text on the flagship Sonic-3.6 text-to-speech model — roughly 1 credit per character — which the pricing page translates into included minutes: roughly 27 TTS minutes/month on Free, ~133 on Pro, ~1,667 on Startup, and ~10,667 on Scale. Ink-2 speech-to-text bills 3 credits per second of audio instead. Localizing a voice costs 225 credits per added voice accent. Cartesia sunset its Voice Changer API on 2026-08-20, and the Pro Voice Clone per-character multiplier (last published at roughly 1.5 credits per character) is no longer shown on Cartesia's live pricing or docs pages as of this revalidation.
- Does Cartesia have a free tier and is a credit card required?
- Yes. The free tier provides 20,000 credits per month with no credit card required at signup, plus a $1/month prepaid balance for voice agents. Free-tier users get the Sonic-3.6 text-to-speech and Ink-2 speech-to-text models, unlimited workspace seats, and one voice agent slot.
- How is Cartesia's voice-agent pricing different from per-second TTS?
- Cartesia's Line voice agents are billed at a flat $0.06 per minute of call duration on every plan — Free, Pro, Startup, and Scale all pay the same rate — plus $0.014 per minute of telephony when you use a Cartesia-provided phone number. Agents draw from a prepaid dollar balance ($1/mo on Free up to $299/mo on Scale) rather than a bundle of included minutes.
- What is the difference between instant and professional voice cloning?
- Instant voice cloning is included from the $5 Pro tier and produces a usable clone from a short reference clip. Professional (pro) voice cloning is introduced at the $49 Startup tier. Localizing a voice into another language costs 225 credits per added voice accent — the pricing page described this as a "one-time cost" until an August 2026 copy update, though the credit amount itself has not changed — rather than a flat dollar setup fee.
- Does Cartesia offer on-premise deployment?
- Yes, through the Enterprise tier. Cartesia is one of the few voice AI vendors that advertises on-premise or virtual-private-cloud deployment of its Sonic and Ink models, targeting healthcare (BAAs), financial services, and compliance-driven customers with strict data-residency requirements. Enterprise pricing is custom and adds DPAs/BAAs, SSO, and security reviews.