AI Summary
About
Sarvam AI is a Bengaluru-based foundation-model company building a full-stack sovereign AI platform for India: open-weight Indic large language models plus speech (ASR/TTS), translation, transliteration, and document-digitisation APIs tuned for 22 Indian languages. It sells to Indian developers, enterprises, and the public sector — and its pricing reflects that focus, denominated entirely in Indian rupees rather than US dollars. Founded in August 2023 by Vivek Raghavan and Pratyush Kumar (both veterans of AI4Bharat and the EkStep/Aadhaar ecosystem; the legal entity is Axonwise Private Limited), Sarvam raised about $41M in December 2023 led by Lightspeed with Peak XV and Khosla — at the time the largest early-stage Indian AI round — and later a ~$200M Series B at an estimated ~$1.2B valuation, with reports of a further raise toward a ~$1.5B valuation.
What sets Sarvam apart from the rest of this foundation-model cluster is its government anchoring. In 2025 it was selected under India’s IndiaAI Mission to build the country’s first homegrown sovereign foundation model, backed by a reported ₹99 crore ($11M) compute subsidy and access to 4,096 Nvidia H100 SXM GPUs provisioned through Yotta Data Services. That mandate is the strategic spine of the whole price sheet: Sarvam is building India-trained models on subsidised national infrastructure, then metering hosted inference on those models at rupee-native rates a domestic buyer can budget against.
The model catalog runs from Sarvam-M — a 24B open-weights hybrid launched in May 2025, built on top of Mistral Small and fine-tuned for Indian languages, math, and code (it drew “foreign model in a desi kurta” criticism for that lineage) — to the Sarvam-30B (mixture-of-experts) and Sarvam-105B (activates ~9B params/token, 128K context) models launched in February 2026 and trained from scratch in Bengaluru. All are open-weighted on Hugging Face, so buyers can self-host the same models they could call over the API — the open-weight hedge familiar from Mistral AI, here wrapped in a sovereign-AI mission rather than a European one.
Pricing summary : How Sarvam AI’s pricing model works
Sarvam AI runs a pure usage-based model with a freemium on-ramp, and every meter is priced in Indian rupees. There is no per-seat subscription — you draw down a shared credit balance and each API meters it at a per-unit rate. As of 2026-08-28 the marketing pricing page (sarvam.ai/api-pricing) was rebuilt around a 2-card hero (Starter “Pay as you go” / Enterprise “Custom pricing”) plus a detailed accordion rate table with a five-rung tier ladder — Pay-As-You-Go, Starter, Pro, Growth, Business — applied to some meters. The dimensions are:
- LLM tokens — separate input, cached-input, and output rates per million tokens, varying by model, priced flat regardless of tier. The long-standing ~7× gap between Sarvam’s developer docs and its marketing pricing page is now resolved: both surfaces agree on Sarvam-105B at ₹29.28 in / ₹10.98 cached / ₹73.2 out, and both now list the two beta models — Gemma-4 31B (₹36.6 / ₹13.73 / ₹91.5) and GLM-5.2 (₹128.1 / ₹23.79 / ₹402.6) — that previously appeared only in docs. Sarvam-30B, previously still sold on the marketing page at ₹2.5/₹1.5/₹10 despite being marked deprecated in docs, has been removed from marketing entirely.
- Speech-to-text (Saaras) — billed per audio hour (per second, rounded up): ₹30/hr standard, ₹45/hr with speaker diarization, flat on both surfaces (marketing labels the same rates “Real-time / Streaming / Batch / Batch with diarization”).
- Text-to-speech (Bulbul v3) — ₹30 per 10,000 characters on docs, shown as ₹3.00 per 1,000 characters on the redesigned marketing page — the same rate, different unit presentation.
- Translation & text tools — a new discrepancy. Docs still price Sarvam Translate / Mayura / transliteration flat at ₹20/10K characters and language identification at ₹3.5/10K, unchanged since June. Marketing’s rebuilt page instead shows a single tiered “Doc translation” line at ₹0.005/character (Pay-As-You-Go through Pro) tapering to ₹0.0045 (Growth) and ₹0.004 (Business) — i.e. roughly ₹40–50 per 10K characters, 2–2.5× the docs rate. Document parsing: Digitisation API ₹0.5/page on both surfaces; marketing also now lists a separate Extraction API at ₹1.00/page that does not appear on docs.
- Dubbing — docs price it per second of source media (truncated to whole seconds) × number of target languages: the standard API rate (
editor_flow: false) is ₹40/min on Starter, ₹38/min on Pro, ₹36/min on Enterprise, doubling to ₹80/₹75/₹72/min with the interactive Editor Flow. Marketing’s rebuilt page instead shows a per-second tiered rate — ₹1.67/sec for Pay-As-You-Go through Pro, tapering to ₹1.25/sec (Growth) and ₹1.20/sec (Business) — that does not cleanly map onto either of the docs tables. - Tier structure — the marketing hero now shows only two purchasable cards (Starter / Enterprise); the specific prepay-plus-bonus credit amounts previously shown for Pro (₹10,000 + ₹2,000 bonus) and Business (₹50,000 + ₹12,500 bonus) are no longer published anywhere on the page. The account-level rate-limit ladder is unchanged: Starter 60, Pro 200, Business 1,000, Enterprise custom requests/minute — a 4-tier ladder that does not include the rate table’s separate “Growth” tier.
What makes this different: Sarvam prices its entire stack in rupees with no USD card at all — a deliberately sovereign, geo-native price built for the India market — and serves open-weight models trained on government-subsidised national compute, so the per-token rate is for hosted convenience, not the model itself.
Pricing by product
LLM API — chat & reasoning (per 1,000,000 tokens, INR)
As of this 2026-08-28 capture, Sarvam’s developer docs and marketing pricing page agree on LLM pricing — resolving the ~7× gap this page tracked from 2026-08-14 through 2026-08-16 (see Pricing evolution for that history). Both surfaces now show:
| Model | Input /1M | Cached input /1M | Output /1M | Key mechanics |
|---|---|---|---|---|
Sarvam-105B (sarvam-105b) | ₹29.28 | ₹10.98 | ₹73.2 | ~9B active params/token, 128K context, 22 Indian languages |
Sarvam 105B Chat (sarvam-105b-conversations) | ₹29.28 | ₹10.98 | ₹73.2 | Conversational SKU, priced identically |
Gemma-4 31B (gemma4) | ₹36.6 | ₹13.73 | ₹91.5 | Beta; now listed on both docs and marketing |
GLM 5.2 (glm5.2) | ₹128.1 | ₹23.79 | ₹402.6 | Beta; reasoning tokens billed as output. Now listed on both surfaces |
Sarvam-30B — previously still sold on the marketing page at ₹2.5 / ₹1.5 / ₹10 despite being marked “Deprecated” in docs — has been removed from the marketing pricing table entirely; it no longer appears on either surface’s price sheet.
Marketing’s rebuilt LLM table additionally lists a third row, “Sarvam 105B Conversations,” priced identically to “Sarvam 105B Chat” (₹29.28 / ₹10.98 / ₹73.2); docs shows only two rows for this model (sarvam-105b and sarvam-105b-conversations), so this may be the same SKU listed under two labels rather than a distinct model.
Speech APIs (INR)
| Service | Price | Key mechanics |
|---|---|---|
| Saaras speech-to-text | ₹30 / hour | Billed per second, rounded up |
| STT + speaker diarization | ₹45 / hour | Adds speaker labels |
| STT + translation | ₹30 / hour | Transcribe + translate, no surcharge over base |
| STT + translation + diarization | ₹45 / hour | Full pipeline |
| Bulbul text-to-speech v3 | ₹30 / 10K chars | Latest voices |
Marketing’s redesigned page relabels these as “Real-time / Streaming / Batch / Batch with diarization” (Speech to text) and “Real-time / Streaming” at ₹3.00 per 1,000 characters (Text to speech) — the same rates, different service names and a different unit denominator (per 1,000 vs per 10,000 characters). Bulbul v2 (previously ₹15/10K chars) is not listed on either surface in this capture.
Text & document APIs (INR)
Docs and marketing now disagree on translate/transliterate pricing and document parsing — a gap first observed in this 2026-08-28 capture, opening the same day the LLM gap above closed.
docs.sarvam.ai — flat rate, unchanged since 2026-06-11:
| Service | Price | Key mechanics |
|---|---|---|
| Sarvam Translate V1 | ₹20 / 10K chars | Indic translation |
| Translate Mayura V1 | ₹20 / 10K chars | Translation model |
| Transliterate | ₹20 / 10K chars | Script conversion |
| Language identification | ₹3.5 / 10K chars | Cheapest meter on the sheet |
| Doc digitisation (Sarvam Vision) | ₹0.5 / page | Max 10 pages per job |
sarvam.ai/api-pricing — new tiered structure, first seen 2026-08-28:
| Service | Pay-as-you-go | Starter | Pro | Growth | Business | Unit |
|---|---|---|---|---|---|---|
| Doc translation (translation & transliteration) | ₹0.005 | ₹0.005 | ₹0.005 | ₹0.0045 | ₹0.004 | per character |
| Digitisation API | ₹0.50 | ₹0.50 | ₹0.50 | ₹0.50 | ₹0.50 | per page |
| Extraction API (new) | ₹1.00 | ₹1.00 | ₹1.00 | ₹1.00 | ₹1.00 | per page |
Marketing’s “Doc translation” rate (₹0.004–₹0.005/character, i.e. roughly ₹40–50 per 10,000 characters) runs 2–2.5× above the docs’ flat ₹20/10K-character rate for the same-sounding capability (translation + transliteration); it isn’t clear from either page whether this is the same product renamed or a genuinely different SKU. Digitisation API pricing (₹0.5/page) is flat and matches docs exactly. The Extraction API (₹1.00/page) does not appear anywhere on the docs pricing page as of this capture.
Dubbing (INR)
docs.sarvam.ai — per minute of source audio, unchanged since 2026-07-30:
| Plan | Dubbing API (editor_flow: false) | Editor Flow (editor_flow: true) |
|---|---|---|
| Starter | ₹40/min | ₹80/min |
| Pro | ₹38/min | ₹75/min |
| Enterprise | ₹36/min | ₹72/min |
sarvam.ai/api-pricing — per second of source audio, new tiered table first seen 2026-08-28:
| Pay-as-you-go | Starter | Pro | Growth | Business |
|---|---|---|---|---|
| ₹1.67/sec | ₹1.67/sec | ₹1.67/sec | ₹1.25/sec | ₹1.20/sec |
Converted to a per-minute rate (×60): Pay-as-you-go/Starter/Pro ≈ ₹100/min, Growth ≈ ₹75/min, Business ≈ ₹72/min. The Growth and Business figures land close to docs’ Editor-Flow Pro (₹75/min) and Enterprise (₹72/min) rates, but the Pay-as-you-go/Starter/Pro marketing figures (≈₹100/min) don’t match any published docs rate — neither the standard table (₹36–40/min) nor the Editor Flow table (₹72–80/min). Which docs tier, if any, a given marketing tier corresponds to is not resolvable from the captured evidence; both tables are live and unreconciled.
Billed per second of source media, truncated to whole seconds, then multiplied by the number of target languages on both surfaces — a 90-second file dubbed into Hindi and Tamil on Starter at the docs’ default API rate costs roughly ₹120 (90s × 2 languages × ₹40/60 per second). Docs’ editor_flow: false is the default for standard API integrations; setting editor_flow: true costs exactly double and switches to the interactive Creator Studio editor workflow, which also suppresses auto-export.
Sales motions across products: PLG / self-serve for every API and the Starter tier; sales-led only for Enterprise (“Custom pricing”) sovereign/on-prem deployments with custom rate limits, SLAs, and data-residency controls.
Hidden costs : What Sarvam AI users actually pay
Sarvam’s per-unit rates are unusually low and fully public, but the real bill is shaped by three things the headline rate doesn’t show: the output-token premium on the LLM, the fact that speech and text meters are denominated differently (per hour vs per 10K chars), and the rate-limit ceiling that effectively forces a prepaid upgrade for production traffic. Two archetypes show how the total assembles.
Archetype 1 — a Hindi voice-assistant startup. Transcribing 2,000 hours of call audio a month with Saaras (with diarization), generating 50M characters of Bulbul v3 voice replies, and routing 40M input plus 10M output tokens a month through Sarvam-30B for the conversation logic.
| Line item | Monthly cost |
|---|---|
| Saaras STT (diarized) — 2,000 hrs @ ₹45/hr | ₹90,000 |
| Bulbul v3 TTS — 50M chars @ ₹30 / 10K | ₹1,50,000 |
| Sarvam-30B input — 40M tok @ ₹2.5/1M | ₹100 |
| Sarvam-30B output — 10M tok @ ₹10/1M | ₹100 |
| Estimated total |
The lesson: for a voice product the speech meters dominate, not the LLM. Tokens are almost free at these rates — the bill is overwhelmingly TTS characters and ASR hours, so the value metric to optimize is audio minutes and spoken characters, not prompt size. A product team must read all three meters together because they scale on completely different units.
Archetype 2 — a translation pipeline at production scale. Translating 500M characters of catalog/content a month via Sarvam Translate, with language ID on each item.
| Line item | Monthly cost |
|---|---|
| Sarvam Translate — 500M chars @ ₹20 / 10K | ₹10,00,000 |
| Language ID — 500M chars @ ₹3.5 / 10K | ₹1,75,000 |
| Estimated total |
Here the surprise is the rate limit, not the rupees: 500M chars/month at production cadence will blow past the Starter 60 req/min and even the Pro 200 req/min ceiling, so the real cost of “scale” is moving to the Business plan (₹50,000 prepay, 1,000 req/min) or an Enterprise quote. The per-character price is cheap; the throughput tier is the gating cost.
Want to estimate your own Sarvam AI bill? Use the Sarvam AI pricing calculator to model your costs across tokens, audio hours, and characters.
Pricing evolution : Sarvam AI pricing history and changes
Sarvam’s pricing followed its model roadmap. There was no public API price until the first hosted model shipped; the rupee-native usage sheet appeared with Sarvam-M in May 2025 and broadened as the from-scratch Sarvam-30B/105B models landed in early 2026. The milestones below are reconstructed from primary announcements and contemporaneous press; quarter-level cadence will be tightened with archived snapshots on a later pass.
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2023 Q4 | 0 | 0 | ~$41M seed + Series A; pre-product, no public pricing |
| 2025 Q2 | 1 | 1 | 2025-05 Sarvam-M ships; first public INR usage API + free credits |
| 2025 Q2 | 0 | 0 | 2025-04/05 Selected under IndiaAI Mission for the sovereign model |
| 2026 Q1 | 0 | 1 | 2026-02-18 Sarvam-30B + Sarvam-105B (from-scratch) become the priced API models |
| 2026 Q2 | 0 | 0 | Live INR price sheet verified across LLM, speech, and text meters |
| 2026 Q3 | 0 | 0 | 2026-07-23 Prepaid credits repackaged — Pro/Business bonuses raised, “Most Popular” moves to Business; per-unit rates held |
| 2026 Q3 | 1 | 2 | 2026-08-14 Docs pricing page reprices Sarvam-105B ~7x higher (₹29.28/₹10.98/₹73.2 vs the marketing page’s unchanged ₹4/₹2.5/₹16) and adds two new beta models, Gemma-4 31B and GLM-5.2 — unreconciled with marketing as of this writing |
| 2026 Q3 | 0 | 1 | 2026-08-15 Dubbing API pricing (₹36–₹80/min) surfaces on the docs pricing page — live since at least 2026-07-30 but undocumented here until now |
| 2026 Q3 | 2 | 1 | 2026-08-28 Marketing pricing page rebuilt: Sarvam-105B repriced ~7x to match docs (closing the gap opened 2026-08-14) and Sarvam-30B delisted; Translate and Dubbing reprice on marketing instead, opening a new, unreconciled gap with docs; a marketing-only Extraction API SKU appears |
Tracked range: 2023 Q4–2026 Q3. Quarters not listed had no publicly announced price or SKU change. Per-snapshot price reconstruction is a later pass; the api-pricing page now resolves (it previously 404’d), and prices read from api-pricing + docs.
Notable changes
- 2023-12 — ~$41M seed + Series A led by Lightspeed (Peak XV, Khosla) — largest early-stage Indian AI round at the time. No public pricing yet (TechCrunch).
- 2025-04/05 — Selected under India’s IndiaAI Mission to build the sovereign foundation model; reported
₹99 crore ($11M) compute subsidy + 4,096 H100 GPUs via Yotta (Inc42). - 2025-05-23 — Sarvam-M (24B open-weights, built on Mistral Small) launches with a public, INR-denominated usage API and free starter credits — the first priced surface.
- 2026-02-18 — Sarvam-30B + Sarvam-105B (from-scratch, fully domestic) launch and become the models behind the paid per-token API (Sarvam-30B ₹2.5/₹10, Sarvam-105B ₹4/₹16 per 1M) (TechCrunch).
- 2026 — ~$200M Series B at an estimated ~$1.2B valuation (Peak XV, Lightspeed), with reports of a further raise toward ~$1.5B.
- 2026-07-23 — Prepaid credits repackaged; per-unit rates held. Every API meter stayed flat (LLM ₹2.5–₹16/1M, STT ₹30–45/hr, TTS ₹15–30/10K, translate ₹20/10K, Vision ₹0.5/page), but the credit ladder that gates rate limits was restructured: the Pro bonus doubled (₹1,000→₹2,000; 11,000→12,000 credits) and the Business bonus rose ~67% (₹7,500→₹12,500; 57,500→62,500 credits), deepening the effective prepay discount to roughly 20% on Pro and 25% on Business. The “Most Popular” badge moved from Pro (₹10,000) to the ₹50,000 Business tier and Business support was trimmed from Slack + dedicated engineer to email — a coherent move to pull production buyers up the ladder while lowering the cost to serve that tier. Starter now shows ₹300 bonus credits; the api-pricing page now resolves (previously 404); Sarvam-M is marked deprecated and Bulbul v3 is flagged as beta pricing.
- 2026-08-14 — Sarvam-105B repriced ~7x on the docs sheet, plus two new beta models. The docs pricing page reprices Sarvam-105B to ₹29.28 input / ₹10.98 cached / ₹73.2 output per 1M tokens — roughly 7x the ₹4/₹2.5/₹16 the same page showed weeks earlier and that the marketing api-pricing page (this page’s
sourceUrl) still advertises unchanged. The docs page also adds two beta-priced models absent from marketing entirely: Gemma-4 31B and GLM-5.2. - 2026-08-16 — Adjudicated: the docs sheet is operative, the marketing sheet is stale. A live re-read of both surfaces found them still disagreeing, with no changelog entry on either side. The marketing page was determined to be the un-maintained one because it still sells the docs-deprecated Sarvam-30B, because its own prose block contradicts its own plan cards on free and bonus credits (quoting the pre-2026-07-23 packaging), and because every SKU shipped since July has landed on docs and never reached marketing. The docs rate card is additionally self-consistent (output = 2.5x input, cached = 0.375x input across both Sarvam-105B and Gemma-4 31B), which a mis-keyed figure would not be. The Facts tables above now lead with the docs rates; the marketing rates are retained as the contradicting surface rather than deleted, because they remain live and quotable.
- 2026-08-15 — A Dubbing API is documented for the first time, three weeks after it went live. The docs pricing page carries a full per-minute Dubbing price table — ₹40/₹38/₹36 per minute (Starter/Pro/Enterprise) billed per second of source audio times target-language count, doubling to ₹80/₹75/₹72/min when the request-level
editor_flow: trueflag switches the job into the Creator Studio editor workflow. Capture evidence shows this pricing was already live on 2026-07-30; it still doesn’t appear on the sarvam.ai marketing pricing page. - 2026-08-28 — Marketing pricing page rebuilt: the LLM gap closes, a new one opens. Sarvam rebuilt sarvam.ai/api-pricing around a 2-card Starter/Enterprise hero plus a 5-tier (Pay-As-You-Go/Starter/Pro/Growth/Business) rate table. Marketing’s Sarvam-105B price jumped to match docs exactly (₹29.28/₹10.98/₹73.2), closing the ~7x gap open since 2026-08-14, and Sarvam-30B — still sold on marketing past its own deprecation — was dropped from the sheet entirely. But the same rebuild opened a new gap the same day: marketing’s tiered “Doc translation” rate (₹40–50/10K chars) now runs 2–2.5x above docs’ unchanged flat ₹20/10K rate, and marketing’s new per-second Dubbing rate doesn’t cleanly map onto docs’ per-minute table. A marketing-only Extraction API (₹1.00/page) appeared, and the specific Pro/Business bonus-credit figures were dropped from the page entirely.
Two surfaces, one widening gap
Read together, the August events are not two isolated data points but one pattern: Sarvam is running its docs pricing page and its marketing pricing page as independently maintained sources of truth that are not reconciled with each other. The Dubbing table sat live and billing on docs.sarvam.ai for at least three weeks (2026-07-30 to 2026-08-15) before anyone outside the company would have found it, since it was never linked from the marketing page or the plan cards. In parallel, the same docs page has shown a Sarvam-105B rate roughly 7x higher than the marketing page’s advertised figure since at least 2026-08-14, with no correction, changelog note, or reconciliation on either side.
Which one is real? A live re-read of both surfaces on 2026-08-16 found them still disagreeing — and settled which to believe. The marketing page is the stale one, on three independent tells. It still sells Sarvam-30B, a model its own docs mark “(Deprecated)” and drop from the price table entirely. It contradicts itself on a single screen: a prose block above the plan cards claims “₹1,000 in free credits”, “Sarvam 105B (Chat LLM): Free per token”, and Pro/Business bonuses of ₹1,000/₹7,500 — figures the cards immediately below it supersede with ₹300, ₹2,000 and ₹12,500, exactly the values set in the 2026-07-23 repackaging. And every SKU Sarvam has actually shipped since July — Document Digitization, Dubbing, Gemma-4 31B, GLM-5.2 — appeared on docs and has still never reached marketing. The docs rate card is also internally coherent in a way a typo would not be: output is exactly 2.5× input and cached exactly 0.375× input across both Sarvam-105B and Gemma-4 31B, whereas the marketing sheet uses a different 4× ratio — two deliberate rate cards from different eras, not one mis-keyed number.
So the divergence is better read not as “Sarvam has two prices” but as “Sarvam repriced its flagship model roughly 7× and never updated the page most buyers will read.” That is a materially worse story than a stale footnote: the page linked from the company’s own site navigation currently advertises a rate a customer cannot get, on a model that no longer exists, to exactly the sovereignty-conscious public-sector and enterprise buyers the IndiaAI mandate is meant to attract. It invites the question of which number a contract should cite — and, if a buyer signed against the marketing figure, whose error it is. The fix is not more pricing pages; it’s a single canonical price feed both surfaces render from, which is precisely the discipline usage-based platforms need once they operate more than one customer-facing price surface.
Update, 2026-08-28: the LLM gap closed — the hard way. Twelve days after that adjudication, Sarvam rebuilt the marketing pricing page and resolved the Sarvam-105B discrepancy by raising the marketing figure to match docs, not the reverse: the price a buyer sees on sarvam.ai jumped roughly 7× overnight, from ₹4/₹2.5/₹16 to ₹29.28/₹10.98/₹73.2 per 1M tokens, and Sarvam-30B — the cheap model marketing had kept selling past its own deprecation — was removed from the sheet entirely. That confirms the 2026-08-16 read: the docs number had been the real one all along, and marketing had simply been stale. But the same rebuild immediately reproduced the pattern it fixed: the new “Doc translation” rate on marketing (₹40–50/10K characters, tiered) now runs 2–2.5× above the unchanged docs rate (flat ₹20/10K), and a new per-second Dubbing rate on marketing doesn’t map cleanly onto the docs’ per-minute table at all. Reconciling one gap, in the same release, opened two more — evidence the root cause was never a single stale number so much as a process gap: whichever surface a rate table gets edited on, it does not appear to be checked against the other before shipping.
From fine-tune to from-scratch, in detail
The most consequential shift in Sarvam’s short history is not a price change but a provenance change that the price sheet now rests on. Sarvam-M (May 2025) was a 24B open-weights fine-tune built on top of Mistral Small — capable on Indic benchmarks but, critics argued, “a foreign model in a desi kurta.” Under the IndiaAI sovereign mandate, the February 2026 Sarvam-30B and Sarvam-105B were trained from scratch in Bengaluru on subsidised national H100 compute. The pricing implication is sovereignty-as-credibility: the same rupee-native per-token rates now buy inference on a genuinely India-built model, which is the exact value proposition a public-sector or sovereignty-conscious enterprise buyer is paying the platform to stand behind.
What’s unique : Sarvam AI’s distinctive pricing mechanics
1. Sovereign, rupee-native pricing — no USD card at all. Almost every foundation-model lab in this corpus prices in dollars (and many geo-lock a rupee or euro view behind a toggle). Sarvam does the opposite: its entire price sheet is denominated in Indian rupees with no USD option, because the buyer it is built for — Indian developers, enterprises, and the government — budgets in rupees. That is not a cosmetic choice; it is the monetization expression of an IndiaAI-Mission sovereign mandate. The pricing is the positioning: a national AI stack priced for the nation that subsidised it.
2. Four differently-denominated meters in one stack. Sarvam bills LLM inference per million tokens, speech per audio hour, text/TTS per 10,000 characters, and — as of the Dubbing API surfacing in August 2026 — video/audio dubbing per minute of source media multiplied by target-language count. That’s four distinct units in a single platform. For a multimodal Indic product (voice assistant, dubbing, doc pipeline) the cost driver shifts surface to surface: tokens are nearly free, but audio hours, spoken characters, and now dubbed minutes-times-languages dominate. Buyers have to model each meter on its own unit, which makes Sarvam’s bill behave very differently from a token-only lab like OpenAI.
3. Open weights on subsidised national compute, then metered hosting. Like Mistral AI, Sarvam open-weights its models (Sarvam-M, 30B, 105B) so buyers can self-host — but the models were trained on government-subsidised H100s under the IndiaAI Mission. So the per-token rate isn’t for the model (which you can download free) and isn’t even fully for the compute (which was subsidised); it’s for managed inference plus the sovereignty assurance, a structurally cheaper-to-produce inference product than a privately-funded lab’s.
4. A request-level boolean, not a plan tier, doubles the price. The Dubbing API’s editor_flow parameter is a per-call flag rather than a plan upgrade: leaving it false bills the standard per-minute rate, and setting it true exactly doubles the charge and switches the job into the interactive Creator Studio workflow. Most usage-based vendors move price by plan or volume tier; Sarvam demonstrates a third lever — a single API parameter that reprices an individual request in real time, a preview of how granular feature-flagged usage pricing can get once the meter is a request rather than a subscription.
5. Selective volume tiering — only the meters facing price competition get cheaper at scale. The 2026-08-28 marketing redesign introduced a five-rung tier ladder (Pay-As-You-Go → Starter → Pro → Growth → Business), but it only discounts two meter families — Translate/transliteration and Dubbing step down per unit as you climb tiers, while LLM tokens, speech, and document APIs stay flat at every tier. That’s a legible signal about where Sarvam feels price pressure: the meters getting volume incentives are the commodity-adjacent ones a generic translation or dubbing vendor could undercut, not the LLM inference where the sovereign, from-scratch-trained model already does the differentiating.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Fully public, rupee-native per-unit rates across LLM, speech, and text — no “contact sales” wall for any meter | INR-only sheet with no USD card adds friction for global buyers who must convert and watch FX |
| Even at its corrected, 7x-higher rate, Sarvam-105B input (₹29.28/1M, ~$0.34) undercuts most Western frontier APIs | Output-token premium on the LLM (₹73.2 vs ₹29.28 input on 105B, 2.5×) favors short-answer workloads; the cheaper Sarvam-30B tier that once softened this is gone as of 2026-08-28 |
| Open weights on Hugging Face let buyers self-host the same models — a credible lock-in hedge | Speech and text use different meters (per hour vs per 10K chars), making blended cost harder to predict |
| Government-anchored sovereign-AI mandate + subsidised compute keep inference cheap and credible for Indian buyers | Rate limits (60/200/1,000 req/min) gate throughput, so production scale forces a prepaid upgrade |
| Credits never expire and roll over indefinitely — no use-it-or-lose-it prepaid trap | Enterprise sovereign/on-prem deployments are fully sales-gated with no public floor price |
| Free starter credits + a free open-weight path make the on-ramp genuinely zero-cost | Early model (Sarvam-M) was a Mistral-Small fine-tune, drawing “desi kurta” provenance criticism |
| Once found, Dubbing pricing is mechanically simple on docs — a flat per-minute rate times language count, with a single boolean flag for the editor workflow | The docs-vs-marketing gap doesn’t stay fixed: the ~7x Sarvam-105B discrepancy (open since 2026-08-14) closed on 2026-08-28 when marketing repriced to match docs, but the same rebuild opened new gaps on Translate (marketing 2–2.5x docs) and Dubbing (marketing’s per-second tiers don’t map to docs’ per-minute table) — reconciliation fixes one number and creates the next |
| Universal credits still apply across every API with no monthly commitment, even after the 2026-08-28 hero redesign | That same redesign removed the specific Pro (₹10,000 + ₹2,000 bonus) and Business (₹50,000 + ₹12,500 bonus) figures from the page entirely — a buyer can no longer see what a mid-tier plan costs without reconstructing it from the separate rate-limit ladder and per-unit rate table |
Billing UX : Sarvam AI billing controls and transparency
- Pay-as-you-go credit wallet, redesigned 2026-08-28 — the marketing hero now shows only two purchasable cards, Starter (“Pay as you go,” add credits as needed) and Enterprise (“Custom pricing,” request a plan); the specific prepay-plus-bonus amounts previously shown for Pro (₹10,000 + ₹2,000 bonus) and Business (₹50,000 + ₹12,500 bonus) are no longer displayed anywhere on the page. Credits remain universal across every API with no monthly seat commitment; docs still states every new user receives ₹100 in free credits.
- Non-expiring, rolling credits — credits “never expire and roll over indefinitely,” so unused balance is never forfeited — a buyer-friendly contrast to the typical expiring-credit model.
- Per-second / per-character billing on metered APIs — speech-to-text is billed per second (rounded up); translation, transliteration, and dubbing are billed per character or per second rather than in whole-unit blocks, so short usage isn’t over-charged.
- A new usage-tier ladder on Translate & Dubbing only — the redesigned detailed pricing table introduces five rungs (Pay-As-You-Go, Starter, Pro, Growth, Business) that step the per-unit rate down as you move up tiers, but only for the Translate and Dubbing meters; LLM, Speech, and Document AI rates stay flat regardless of tier. This 5-rung ladder is separate from — and doesn’t match — the account-level rate-limit ladder below, which still has four rungs (no “Growth”).
- Rate-limit tiers as a separate upgrade lever — the account-level requests/minute ceiling still steps Starter 60 → Pro 200 → Business 1,000 → Enterprise custom, unchanged from prior captures; the “Most Popular” badge that previously marked the ₹50,000 Business card is gone from the redesigned 2-card hero.
- Free self-host path — open weights on Hugging Face let cost-sensitive teams run inference themselves and skip the API meter entirely for non-managed workloads.
editor_flowbilling flag on Dubbing (docs) — a request-level parameter, not a plan setting: Dubbing API calls default toeditor_flow: false(standard per-minute rate, ₹40/min on Starter per docs); flipping it toeditor_flow: truedoubles the rate to ₹80/min on Starter and switches to the interactive Creator Studio editor workflow, which also suppresses auto-export — a per-call price/behavior toggle rather than a plan upgrade.- Plan-specific support tiers are no longer disclosed on marketing — the redesigned 2-card hero lists account-level controls (rate limits, concurrency, volume discounts, SLAs) for Starter and Enterprise only; the prior page’s per-plan support detail for Pro/Business (community vs. email vs. Slack + Solutions Engineer) is not shown on the current page.
Strategic wins : Why Sarvam AI’s pricing decisions worked
1. Pricing the nation’s stack in the nation’s currency
By denominating the entire sheet in rupees with no USD card, Sarvam turned a sovereign-AI mandate into a pricing strategy. For the Indian developer or public-sector buyer it is competing for, a rupee-native rate removes FX friction and signals “built for you.” It is the monetization expression of the IndiaAI Mission selection — the pricing reinforces the positioning rather than fighting it. See usage-based pricing strategy for why aligning the meter to the buyer’s mental model wins.
2. Subsidised compute funds aggressive unit economics
Because the models were trained on government-subsidised H100 compute under the IndiaAI Mission, Sarvam can price hosted inference well under privately-funded frontier labs even at its corrected rate — Sarvam-105B input runs ₹29.28/1M (~$0.34 at ₹85/$), which is roughly 7x what marketing advertised before the 2026-08-28 rebuild but still well below Western frontier list prices. Sarvam-30B, the cheaper illustrative model previously quoted here at ~$0.03/1M input, has been discontinued and removed from the price sheet entirely as part of that same rebuild, so the moat now rests on the flagship model’s rate holding up on its own rather than on a bargain second tier. This mirrors the shift away from rigid per-seat licensing toward cost-following usage rates.
3. Open weights plus non-expiring credits lower the on-ramp to zero
A free open-weight self-host path and free starter credits and credits that never expire combine into an unusually low-risk on-ramp. A developer can prototype free, self-host if they prefer, and never lose prepaid balance — which de-risks adoption in a price-sensitive market. Choosing a durable usage metric and pairing it with a forgiving credit model is what makes that on-ramp stick.
Areas to improve : Gaps in Sarvam AI’s pricing approach
1. Offer a USD view for global buyers
The INR-only sheet is perfect for India but adds conversion friction for the diaspora developers, NRIs, and global teams who want Indic-language inference. A USD toggle (as most peers offer) would widen the addressable market without diluting the sovereign positioning — the rupee can stay the default. The absence today risks reading as “domestic-only,” which understates the models’ reach.
2. Make the blended multi-meter cost legible up front
Because LLM, speech, and text bill on three different units, a multimodal product’s total is hard to forecast from the price sheet alone. A worked “estimated cost per voice session” or per-document calculator surfaced in the dashboard would prevent the bill-shock and unpredictability that mixed meters invite, and help buyers self-select a prepaid tier confidently.
3. Expose an Enterprise / sovereign-deployment anchor
On-prem and sovereign deployments are fully sales-gated with no public starting point. Given that the IndiaAI mandate makes public-sector and regulated buyers the core market, a published floor or a worked deployment example would shorten procurement cycles for exactly the institutional buyers Sarvam is positioned to win. Compare how other AI companies stage enterprise transparency.
4. Reconcile pricing across marketing and docs surfaces — durably, not just reactively
Three August 2026 events point at the same underlying gap, and closing it once didn’t make it stay closed. A full Dubbing price table sat live on docs.sarvam.ai for roughly three weeks (2026-07-30 to 2026-08-15) with no link from the marketing pricing page; the docs page quoted a Sarvam-105B rate roughly 7x higher than marketing from 2026-08-14 until 2026-08-28, when the marketing rebuild finally matched it; and that same rebuild immediately opened two new discrepancies, on Translate (marketing running 2–2.5x above docs) and Dubbing (marketing’s new per-second rate doesn’t map onto docs’ per-minute table). That pattern — fix one gap, open the next, in the same release — is evidence the correction was a one-off patch rather than a process change. For a sovereignty-focused platform courting regulated and public-sector buyers, a price sheet that keeps disagreeing with itself undermines exactly the trust the IndiaAI positioning is built to establish. A single canonical price feed driving both surfaces, with every rate change and new SKU landing on both simultaneously and validated against each other pre-launch, is the only fix that would actually hold.
5. Restore self-serve visibility into what a mid-tier plan actually costs
The 2026-08-28 redesign cut the marketing hero from four named plan cards to two (Starter / Enterprise) and, with it, dropped the specific numbers — Pro’s ₹10,000 prepay + ₹2,000 bonus, Business’s ₹50,000 + ₹12,500 bonus — that previously let a buyer see exactly what a mid-volume commitment cost without contacting sales. The account-level rate-limit ladder (Starter 60 → Pro 200 → Business 1,000 req/min) still exists, but a buyer now has to cross-reference that ladder against the separate five-rung per-unit rate table to reconstruct what Pro or Business would actually cost — work the old four-card page used to do for them. Simplifying the hero is a reasonable UX call; quietly deleting the pricing math that made the old version self-serve is not.
Monetization stack & signals : how Sarvam AI builds & buys its revenue engine
Buys 0 Builds 0 13 open roles
Sarvam's monetization stack is still greenfield: a May 2026 "Product Manager, Monetization & Retention" req wants someone who "has debated Lago vs Orb vs Stripe Billing and has opinions" — the billing/metering layer for its prepaid INR credit wallet is being chosen now, not yet operated. Investment is clearly flowing into the revenue engine: ~5 billing/API-platform and data-engineering roles, plus monetization/Studio GTM and a retention-heavy product cluster.
- Usage billing / metering layer Billing inferred Job post May 2026
“You've debated Lago vs Orb vs Stripe Billing and have opinions”
- Product analytics Analytics inferred Job post 1 Job post 2 May 2026
“Mixpanel, PostHog, Amplitude, whatever the stack is”
- CRM CRM inferred Job post May 2026
“You are highly proficient with Salesforce or HubSpot at an operational level”
- GTM & Strategy, Sarvam Studio MonetizationGrowth seen Jun 3, 2026
- Head of Growth Marketing RetentionGrowth seen Jun 3, 2026
- Product Manager, Monetization & Retention MonetizationRetention seen May 25, 2026
- Product Manager, Growth Retention seen May 25, 2026
- Staff Engineer, API Platform Billing engineering seen May 21, 2026
- Staff Data Engineer Billing engineering seen May 21, 2026
- Backend Engineer - Studio Media Platform Billing engineering seen May 21, 2026
- Engineering — Full Stack AI Engineer Billing engineering seen May 21, 2026
- GTM Manager On-Device AI RevOps seen May 13, 2026
- Engagement Principal, Chanakya RevOps seen May 13, 2026
- Engagement Manager, Chanakya Customer success seen Apr 17, 2026
- Account Manager Customer success seen Apr 17, 2026
- GTM/Product, Vision Customer success seen Apr 17, 2026
- +9 more matched roles
Signals reviewed · derived from public job posts
Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.
Key takeaways
- Price in the buyer’s currency, literally. Sarvam’s rupee-only sheet is a deliberate sovereign signal — the pricing is the positioning. When your strategic edge is “built for this market,” denominating in that market’s currency reinforces it more than any tagline.
- Subsidised inputs become a pricing moat. Government-subsidised compute under the IndiaAI Mission lets Sarvam under-price global APIs on Indic workloads. A structural cost advantage upstream shows up as a durable price advantage downstream.
- Multi-meter stacks need multi-meter thinking. Tokens, audio hours, characters, and now per-minute dubbing scale on different units; for a voice or translation product the LLM is nearly free and the speech/text/dubbing meters dominate. Model every meter on its own unit, not the headline token rate.
- A forgiving credit model lowers adoption risk. Free starter credits, a free self-host path, and credits that never expire combine into a near-zero-risk on-ramp — decisive in a price-sensitive market.
- Provenance is a value metric. Moving from a Mistral-Small fine-tune to from-scratch domestic models is what lets the same rupee rates carry a sovereignty assurance that public-sector buyers actually pay for.
UBP implications
- Currency and geo-denomination are pricing levers, not just settings. Sarvam shows that choosing to price natively in the buyer’s currency — and declining to bolt on a USD card — can be a strategic statement. UBP practitioners targeting a specific market should treat denomination as part of the value proposition, not an afterthought.
- Subsidised or differentiated input costs should flow through to the meter. When an upstream advantage (here, subsidised national compute) lowers your cost to serve, passing it through as a lower unit rate converts an infrastructure edge into a competitive pricing edge — the cleanest way usage pricing turns cost structure into go-to-market.
- Mixed-meter platforms must teach buyers which unit drives the bill — and keep every surface that quotes it in sync, not just once. When a stack bills tokens, audio hours, and characters together, the dominant cost driver shifts by use case, and UBP design has to surface the binding meter per workload or buyers misjudge cost. That discipline extends to publishing: Sarvam’s docs and marketing pages quoted different Sarvam-105B rates from 2026-08-14 until a 2026-08-28 marketing rebuild matched them — and the very same rebuild opened new, unreconciled gaps on Translate and Dubbing the same day. One canonical source, validated on every release, is the only fix that survives a redesign; a one-time reconciliation does not.
Sources
- Sarvam AI API pricing page (accessed 2026-07-23)
- Sarvam AI docs — pricing reference (accessed 2026-07-23)
- Sarvam — India’s full-stack sovereign AI platform (accessed 2026-07-23)
- Sarvam open weights on Hugging Face (accessed 2026-07-23)
- Browse the pricing blueprint corpus
Press coverage of Sarvam’s funding rounds and model launches is cited inline in Pricing evolution where it supports a specific claim.
Bottom line
Sarvam AI prices a full sovereign Indic-AI stack on pure usage, denominated entirely in Indian rupees: an LLM API at ₹29.28/1M input, ₹10.98/1M cached, ₹73.2/1M output (Sarvam-105B — now the single flagship rate on both docs and marketing after Sarvam-30B’s 2026-08-28 discontinuation), Saaras speech-to-text at ₹30–45/hr, Bulbul text-to-speech at ₹30/10K chars, translation at ₹20/10K chars on docs (a disagreeing ₹40–50/10K on the newly-tiered marketing table), and Dubbing at ₹36–80/min on docs (a disagreeing per-second rate on marketing) — all on open-weight models trained from scratch on government-subsidised compute under India’s IndiaAI Mission. The rupee-native sheet is the monetization face of a sovereign-AI mandate, the subsidised compute keeps rates well under global frontier APIs even after the August 2026 correction, and a free open-weight path plus non-expiring credits make the on-ramp nearly free. The main friction is the INR-only view for global buyers, a mixed-meter bill that needs per-unit modeling, a docs-vs-marketing price sheet that closed one gap (Sarvam-105B, resolved 2026-08-28 in docs’ favor) only to open two more the same day (Translate, Dubbing), and a redesigned hero that no longer discloses what a mid-tier Pro or Business plan actually costs.
Want to compare Sarvam AI against other foundation-model providers? See Mistral AI and OpenAI, or browse the full pricing blueprint.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
Marketing pricing page rebuilt: LLM gap with docs resolved, Translate/Dubbing gap opens instead
sarvam.ai/api-pricing was rebuilt around a 2-card Starter/Enterprise hero plus a 5-tier (Pay-As-You-Go/Starter/Pro/Growth/Business) detailed rate table. Marketing's LLM prices now match docs exactly (Sarvam-105B ₹29.28/₹10.98/₹73.2, resolving the ~7x gap tracked since 2026-08-14) and Sarvam-30B was removed from marketing entirely. But a new gap opened on Translate and Dubbing: marketing's tiered 'Doc translation' rate (₹40–50/10K chars) runs 2–2.5x above docs' unchanged flat ₹20/10K rate, and marketing's per-second Dubbing rates don't cleanly map to docs' per-minute editor_flow table. A new Extraction API (₹1.00/page) appeared on marketing only. The old marketing prose block that contradicted its own plan cards is gone.
Undocumented Dubbing API surfaces on docs pricing page (per-minute, editor_flow billing toggle)
Sarvam's docs pricing page carries a Dubbing product not previously reflected in these facts: ₹40/₹38/₹36 per minute (Starter/Pro/Enterprise) for standard API dubbing (editor_flow: false), billed per second of source audio multiplied by the number of target languages, doubling to ₹80/₹75/₹72 per minute with the interactive Editor Flow (editor_flow: true). First appeared in capture evidence on 2026-07-30 but was not written up until this pass; still absent from the sarvam.ai marketing pricing page. All LLM, speech, and text per-unit rates and prepaid plan packaging held steady across both pricing surfaces.
Prepaid credits repackaged; 'Most Popular' moves to Business
All per-unit API rates held (LLM ₹2.5–₹16/1M, STT ₹30–45/hr, TTS ₹15–30/10K, translate ₹20/10K, Vision ₹0.5/page), but the prepaid credit packaging changed: Pro bonus ₹1,000→₹2,000 (11,000→12,000 credits), Business bonus ₹7,500→₹12,500 (57,500→62,500 credits), Starter now shows ₹300 bonus credits, and the 'Most Popular' badge moved from Pro to the ₹50,000 Business tier. The api-pricing page now returns 200 (previously 404). Sarvam-M is marked deprecated; Bulbul v3 flagged as beta pricing.
Live INR price sheet: per-token LLM + per-hour speech + per-char text
Captured live INR pricing across the stack: LLM API Sarvam-30B ₹2.5 in / ₹1.5 cached / ₹10 out and Sarvam-105B ₹4 / ₹2.5 / ₹16 per 1M tokens; Saaras STT ₹30/hr (₹45 with diarization); Bulbul TTS ₹15–30/10K chars; translation/transliteration ₹20/10K, language ID ₹3.5/10K, doc parsing ₹0.5/page; prepaid plans Starter (free) / Pro ₹10,000 (+₹1,000) / Business ₹50,000 (+₹7,500) with 60/200/1,000 req-min limits and non-expiring credits. The naked /pricing path 404s; prices read from api-pricing + docs.
Sarvam-30B + Sarvam-105B (from-scratch, fully domestic) launch
Sarvam releases two open-source models trained from scratch in Bengaluru: Sarvam-30B (mixture-of-experts) and Sarvam-105B (activates ~9B params/token, 128K context, 22 Indian languages) — India's first fully domestically-trained open LLMs under the IndiaAI mandate. These become the models behind the paid per-token API (Sarvam-30B ₹2.5/₹10, Sarvam-105B ₹4/₹16 per 1M). (Source: TechCrunch, Sarvam, 2026-02.)
Sarvam-M (24B open-weights) launches with public API + free credits
Sarvam ships Sarvam-M, a 24B open-weights hybrid model built on top of Mistral Small and fine-tuned for Indian languages, math and code (+86% on romanised GSM-8K). It is served via Sarvam's API, playground and Hugging Face — the first model behind a public, INR-denominated usage API with free starter credits. Critics call it 'a foreign model in a desi kurta.' (Source: Sarvam blog/X, Entrepreneur, 2025-05.)
Selected under India's IndiaAI Mission to build the sovereign model
The Government of India selects Sarvam under the IndiaAI Mission to build the country's first homegrown sovereign foundation model, backed by a reported ~₹99 crore (~$11M) compute subsidy and 4,096 Nvidia H100 SXM GPUs provisioned via Yotta Data Services (a 100% compute subsidy reported). This government anchoring shapes the later geo-native INR price sheet. (Source: Inc42, NVIDIA blog, 2025.)
Sarvam raises ~$41M seed + Series A
Five months after its August 2023 founding in Bengaluru, Sarvam (legal entity Axonwise Private Limited) raises about $41M led by Lightspeed with Peak XV and Khosla — at the time the largest early-stage funding for an Indian AI startup. No public API pricing yet; the company is in model-build mode. (Source: TechCrunch, Sarvam blog, 2023-12.)
- · Sarvam's price sheet is denominated entirely in Indian rupees (₹) with no USD card — a deliberately geo-native price for the India market, where Sarvam-30B input runs ₹2.5 per 1M tokens (~$0.03).
- · Sarvam was selected under India's IndiaAI Mission to build the country's sovereign foundation model, backed by a reported ~₹99 crore (~$11M) GPU subsidy and 4,096 Nvidia H100 GPUs via Yotta — a government-anchored sovereign-AI story.
- · Its first hosted model, Sarvam-M (May 2025), was a 24B fine-tune built on top of Mistral Small — drawing 'foreign model in a desi kurta' criticism — before the February 2026 Sarvam-30B/105B models were trained from scratch in Bengaluru.
Questions & answers
- What is Sarvam AI's pricing model?
- Pure usage-based, billed in Indian rupees. The LLM API charges per million tokens — Sarvam-105B at ₹29.28 in / ₹73.2 out, a rate the developer docs and marketing pricing page now agree on as of 2026-08-28. Speech-to-text is ₹30/hr, text-to-speech is ₹30 per 10K characters, and a Dubbing API runs ₹36–80 per minute of source audio on docs (marketing shows a different, higher per-second tiered rate). Translation is ₹20 per 10K characters on docs but a newly-tiered ₹40–50 per 10K characters on the redesigned marketing page — a discrepancy that opened the same week the LLM gap closed. You draw down a shared credit balance that covers every API.
- Does Sarvam AI offer a free tier?
- Yes. As of the 2026-08-28 marketing redesign, the Starter card reads simply 'Pay as you go — start with free credits, then add exactly what you need,' with no specific bonus amount shown (the previous ₹300-bonus figure and the older ₹1,000 prose-block figure are both gone from the page). The developer docs still state every new user receives ₹100 worth of free credits. Starter includes a 60 requests/minute rate limit. Sarvam also open-weights its models (Sarvam-105B and the deprecated Sarvam-30B and Sarvam-M) on Hugging Face for self-hosting.
- How much does the Sarvam LLM API cost per token?
- ₹29.28 per 1M input tokens, ₹10.98 cached, ₹73.2 output for Sarvam-105B — as of 2026-08-28 this rate is confirmed on both Sarvam's developer docs and its marketing pricing page, which previously disagreed by roughly 7x (marketing had advertised ₹4/₹2.5/₹16 for the same model through at least 2026-08-16). Sarvam-30B, previously still sold on marketing despite being marked deprecated in docs, has been removed from the marketing price table entirely. Beta models Gemma-4 31B (₹36.6/₹13.73/₹91.5) and GLM-5.2 (₹128.1/₹23.79/₹402.6) are now listed on both surfaces.
- How does Sarvam price speech and translation APIs?
- Saaras speech-to-text is ₹30/hour (₹45/hour with speaker diarization), billed per second rounded up. Bulbul v3 text-to-speech is ₹30/10K characters. Docs price translation, transliteration and Mayura flat at ₹20/10K characters and language identification at ₹3.5/10K; document parsing (Digitisation API) is ₹0.5/page on both surfaces. As of 2026-08-28 the marketing page instead shows a tiered 'Doc translation' rate of ₹0.004–₹0.005 per character (≈₹40–50/10K chars, roughly 2–2.5x the docs rate) and a new Extraction API at ₹1.00/page that docs doesn't list.
- What is Sarvam's connection to the IndiaAI Mission?
- In 2025 Sarvam was selected under India's government-run IndiaAI Mission to build the country's sovereign foundation model, backed by a reported ~₹99 crore (~$11M) GPU-compute subsidy and access to 4,096 Nvidia H100 GPUs via Yotta. That mandate underpins the from-scratch Sarvam-30B and Sarvam-105B models the paid API now serves.
- Can I self-host Sarvam's models instead of using the API?
- Yes. Sarvam open-weights its models on Hugging Face — Sarvam-M (24B, built on Mistral Small), and the from-scratch Sarvam-30B and Sarvam-105B. You can download and self-host them, and pay Sarvam's per-token API only when you want hosted inference.