AI Summary
About
SambaNova Systems is a Palo Alto AI company founded in 2017 by Stanford professors Kunle Olukotun and Christopher Ré together with former Oracle executive Rodrigo Liang. Rather than building on NVIDIA GPUs, SambaNova designs its own AI silicon — the Reconfigurable Dataflow Unit (RDU), most recently the SN40L and the agentic-inference SN50 — and packages it into full “chips-to-model” systems. The company raised a $676M Series D led by SoftBank Vision Fund 2 in 2021 at a valuation above 5B, pushing total funding past 1B, and announced a further $350M Series E in February 2026 alongside the SN50 and an Intel collaboration. On July 8, 2026 it announced the first close of a $1B financing round at an $11B valuation, led by General Atlantic with Seligman Ventures and T. Rowe Price Associates — with JPMorganChase named as a new customer selecting SambaNova RDUs for on-prem inference.
For pricing purposes, SambaNova is really two businesses. SambaNova Cloud (branded SambaCloud) is a developer-facing, OpenAI-compatible inference API that rents access to open models — Llama, DeepSeek, Qwen-class, gpt-oss, Gemma, MiniMax — billed per million tokens, with a published rate card and a free tier. SambaStack / SambaManaged is the enterprise hardware side: RDU systems and racks sold as sales-quoted contracts with no public price. The throughline is the chip: SambaNova competes less on the cheapest token and more on the fastest token, routinely claiming record tokens-per-second on its own hardware.
For current rates, see SambaNova Cloud pricing. Note the rate card lives on the cloud.sambanova.ai subdomain — the marketing site’s /pricing path returns a 404 because the systems business is sales-only.
Pricing summary : How SambaNova’s pricing model works
SambaNova’s pricing is hybrid, split cleanly by product:
- SambaNova Cloud (inference API) — pure usage-based, billed per 1M tokens with separate input and output rates per model, plus a cached-input rate for models that support prompt caching (MiniMax-M2.7 is the first on the card, at $0.06/1M cached input vs. $0.60 uncached). It has three account tiers: a Free plan ($0, but as of August 2026 requires adding a payment method to purchase pay-as-you-go credits — the earlier no-card free-credit grant was removed), a pay-as-you-go Developer plan, and a subscription-based Enterprise plan with production rate limits and add-ons like BYOC and custom limits. This rate card is fully public.
- RDU systems (SambaStack / SambaManaged / DataScale) — sales-quoted. There is no public price for the hardware, racks, or managed deployments; these are enterprise contracts sold by SambaNova’s go-to-market team.
So the buyer journey is genuinely self-serve at the bottom (sign up, add a payment method, buy credits, call the API) and sales-led at the top (buy or rent RDU capacity), with the per-token API serving as both a product and a demand-generation funnel into the silicon.
What makes this different: Most inference APIs are reselling NVIDIA GPU time and compete on price-per-token. SambaNova runs the same open models on its own RDU silicon and competes on speed-per-token — the rate card is the wrapper, but the pitch is “fastest inference,” not “cheapest.” That makes its per-token prices closer to mid-pack while its differentiation lives in throughput and latency.
Pricing by product
SambaNova Cloud per-1M-token rates, as of August 2026 (USD). The rate card carries a separate Cached Input Tokens column for models that support prompt caching:
| Model | Cached input /1M | Input /1M | Output /1M | Notes |
|---|---|---|---|---|
| gemma-4-31B-it | N/A | $0.38 | $1.15 | Reverted here in August 2026 after a brief July cut to $0.22/$0.59 |
| gpt-oss-120b | N/A | $0.22 | $0.59 | Cheapest on the card; open-weight reasoning |
| MiniMax-M2.7 | $0.06 | $0.60 | $2.40 | Only model with a cached-input rate; high output cost |
| Meta-Llama-3.3-70B-Instruct | N/A | $0.60 | $1.20 | Mainstream workhorse |
| DeepSeek-V3.1 | N/A | $3.00 | $4.50 | Frontier-class, 671B params |
| DeepSeek-V3.2 | N/A | $3.00 | $4.50 | Frontier-class, priciest |
Account tiers (SambaNova Cloud):
| Tier | Price | Included | Key mechanics |
|---|---|---|---|
| Free | $0 | Pay-as-you-go credits, Production models | Requires a payment method to purchase credits — as of August 2026 there’s no automatic free-credit grant |
| Developer | Pay-as-you-go | All Production & Preview models | Standard rate limits, per-token billing; “Add Card to Upgrade” |
| Enterprise | Subscription / custom | Production rate limits, BYOC | Sales-quoted for larger usage |
Sales motions across products: the cloud API Free and Developer tiers are fully self-serve (PLG); Enterprise and all RDU hardware (SambaStack, SambaManaged, DataScale) are sales-led and quoted. There is no public price for the systems business.
Hidden costs : What SambaNova users actually pay
On the cloud side the rate card is clean, but real bills depend on a few things beyond the headline per-token number:
| Line item | Cost |
|---|---|
| Input tokens (e.g. Llama-3.3-70B) | $0.60 per 1M |
| Output tokens (e.g. Llama-3.3-70B) | $1.20 per 1M |
| Reasoning / “thinking” tokens | Billed as output — DeepSeek-V3.2 reasoning at $4.50/1M output adds up fast |
| Free credits | None automatically — as of August 2026 a payment method is required to purchase pay-as-you-go credits |
| Enterprise rate limits / dedicated capacity | Sales-quoted (subscription) |
| RDU systems / SambaStack | Sales-quoted; no public price |
The real cost traps are structural, not line-item. First, output and reasoning tokens dominate — output rates run 2–5x input (MiniMax-M2.7 is $0.60 in but $2.40 out), so chatty or chain-of-thought workloads cost far more than the input-side rate suggests. Second, the DeepSeek-V3.1/V3.2 frontier tier at $3.00/$4.50 is roughly 14x the cheapest model (gpt-oss-120b at $0.22 input), so model choice swings the bill enormously — and that spread isn’t fixed: gemma-4-31B-it briefly matched the $0.22 floor in July 2026 before reverting to $0.38/$1.15 in August, so per-model rates can move without warning. Third, the no-card free trial is gone: as of August 2026 the Free plan requires adding a payment method before you can purchase credits and call the API at all, removing the friction-free evaluation path that used to bridge a slow procurement cycle. And on the systems side, the entire cost is opaque until you talk to sales.
Want to estimate your own SambaNova Cloud bill? Use the SambaNova pricing calculator to model your costs based on model and token volume.
Pricing evolution : SambaNova pricing history and changes
Cadence
| Period | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2024 H2 | Public token rate card launched | SambaNova Cloud (free + pay-as-you-go) | OpenAI-compatible API on RDU |
| 2025 H2 | Per-model rates tracked | Sovereign-AI regional clouds | Argyll, Infercom, OVHcloud, SouthernCrossAI |
| 2026 Q1–Q2 | Rate card spans $0.15–$3.00 input | SN50 RDU; $350M Series E | Newer DeepSeek/gpt-oss/Gemma/MiniMax models added |
| 2026 Q3 | gemma-4-31B-it cut to $0.22/$0.59 (Jul), then reverted to $0.38/$1.15 (Aug); cached-input rate added | $1B financing at $11B valuation | Card trimmed to 6 models; MiniMax-M2.7 gets $0.06 cached input; gemma round-trips within ~3 weeks; Free plan’s no-card $5 credit grant removed Aug 14, now requires a payment method before the first request (token rates untouched) |
Tracked range: 2024–present. The systems/hardware business has never published a public price, so only the cloud rate card is trackable.
Notable changes
- 2024 H2 — SambaNova Cloud launches as a public, OpenAI-compatible inference API with a free developer tier and per-token pay-as-you-go billing, positioned on fastest-token throughput for open models rather than per-GPU-hour rental.
- Late 2025 — Sovereign-AI inference partnerships (UK, Germany, EU, Australia) extend the token-based cloud regionally while keeping the published rate card.
- June 2026 — Rate card spans $0.15/$0.75 (DeepSeek-V3.1-cb) to $3.00/$4.50 (DeepSeek-V3.1/V3.2), with Meta-Llama-3.3-70B at $0.60/$1.20 and gpt-oss-120b at $0.22/$0.59. SN50 RDU and a $350M Series E announced in February 2026.
- July 2026 — Card trimmed to six models and gemma-4-31B-it cut from $0.38/$1.15 to $0.22/$0.59; DeepSeek-V3.1-cb ($0.15/$0.75) and DeepSeek-R1-Distill-Llama-70B ($0.70/$1.40) dropped from the public card. A new Cached Input Tokens column appears, with MiniMax-M2.7 priced at $0.06/1M cached input. On July 8, 2026 SambaNova announced the first close of a $1B round at an $11B valuation (General Atlantic-led), with JPMorganChase named as a new RDU customer.
- August 11, 2026 — gemma-4-31B-it’s price cut is reversed: input/output moves back from $0.22/$0.59 to $0.38/$1.15, matching its pre-July rate. The other five models on the card (MiniMax-M2.7, DeepSeek-V3.1, DeepSeek-V3.2, gpt-oss-120b, Meta-Llama-3.3-70B) are unchanged, leaving gpt-oss-120b as the sole model at the $0.22 price floor.
- August 14, 2026 — SambaNova drops the no-card $5 free-credit grant on SambaNova Cloud’s Free plan. The plans page no longer promises “$5 of Free Credit,” “No credit card required,” or a 30-day expiry window; it now reads “Add a payment method and purchase credits to run your first requests.” Free is still nominally $0 with no subscription fee, but there’s no longer a way to call the API before adding a card. The per-model token rate card is unchanged.
The direction of travel is model proliferation and rate-card cleanup, not flat price moves on the token side, paired with a tightening of the onboarding gate: SambaNova keeps swapping newer open models onto tiered rates, trimming older ones, layering in prompt-caching (cached-input) pricing, and — as the gemma round-trip shows — individual model rates can move in either direction within weeks, so the effective cost depends almost entirely on which model you pick and when you check. The August 14 change is a different axis: it doesn’t move a single price, but it raises the bar to try the API at all, trading a frictionless trial for earlier billing-detail capture.
What’s unique : SambaNova’s distinctive pricing mechanics
1. Speed as the value metric, not price. SambaNova prices per token like everyone else, but the product it’s actually selling is throughput on custom RDU silicon. Its marketing leads with record tokens-per-second, so buyers pay mid-pack token rates for top-tier latency rather than the cheapest possible token.
2. A card-gated free tier on inference, sales-only on hardware. The cloud API still lists a $0 Free plan, but as of August 14, 2026 it requires adding a payment method before you can purchase pay-as-you-go credits and call the API — the earlier no-card $5 grant is gone. RDU systems stay a fully gated, contact-sales motion. The split between self-serve API and enterprise hardware sale is still there within one brand; what changed is that the API side no longer has a zero-friction entry point.
3. Per-model price spread, not per-tier. Instead of bundling tokens into plan tiers, SambaNova lets the model choice set the price: from $0.22 input for gemma-4-31B-it and gpt-oss-120b to $3.00 input for the frontier DeepSeek-V3.1/V3.2 — roughly a 14x spread on the same rate card. The July 2026 refresh narrowed the range by dropping the sub-$0.22 DeepSeek-V3.1-cb, so the cheapest model is now the price floor rather than a distilled outlier.
4. Cached-input pricing enters as a third price axis. As of July 2026 the rate card carries a separate Cached Input Tokens column, with MiniMax-M2.7 billed at $0.06/1M cached versus $0.60/1M uncached — a ~10x discount that rewards repeated-prompt and long-context workloads. It moves SambaNova from a two-number (input/output) meter toward a three-number one, and signals more models will likely gain cached rates.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Public, transparent per-token rate card | Marketing-site /pricing 404s; rate card hidden on cloud subdomain |
| Free plan still carries no subscription fee ($0) | Card required before the first API request — the no-card $5 trial credit was removed Aug 14, 2026 |
| Differentiated on inference speed (custom RDU) | Token rates are mid-pack, not cheapest |
| OpenAI-compatible API, easy migration | Hardware/systems pricing fully opaque (sales-only) |
| Newer open models added quickly | Output/reasoning tokens make bills hard to predict |
Billing UX : SambaNova billing controls and transparency
- Billing controls — Self-serve console issues an API key; as of August 2026 the Free tier requires adding a payment method to purchase pay-as-you-go credits before you can run requests (the earlier no-card free-credit grant is gone), and the Developer tier’s “Add Card to Upgrade” flow unlocks standard rate limits and Preview models. Enterprise moves to subscription-based pricing with production rate limits.
- Usage visibility — Per-token billing with separate input, output and (where supported) cached-input rates is shown on the public pricing page; consumption is metered against credits, then the card. The console carries dedicated Pricing, Billing, Usage, and Commits-and-Credits views.
- Payment options — Self-serve credit-card checkout for Free/Developer; sales-led contracts, invoicing, and BYOC/custom-rate-limit arrangements for Enterprise and all RDU hardware.
Strategic wins : Why SambaNova’s pricing decisions worked
1. Using a $0 token tier as a funnel into custom silicon
Through mid-August 2026, a no-card $5 credit let any developer try RDU-backed inference in minutes; as of August 14, 2026 the Free plan requires a payment method before the first request. The funnel logic survives — a $0 plan is still the cheapest way to reach the API — but SambaNova now trades a bit of top-of-funnel friction for billing-detail capture and abuse-resistance at signup rather than at conversion. See how AI companies structure pricing.
2. Competing on speed instead of racing token prices to zero
By anchoring on fastest-inference rather than cheapest-token, SambaNova avoids the deflationary token price war and justifies mid-pack rates with throughput — a value-metric choice. Related: outcome-based pricing trends.
3. Letting model choice carry the price spread
Rather than rigid plan tiers, SambaNova prices each model independently across a ~14x range (July 2026, after the sub-$0.22 DeepSeek-V3.1-cb was retired), so customers self-select cost/quality without a packaging negotiation. See choosing the right usage metric.
Areas to improve : Gaps in SambaNova’s pricing approach
1. Discoverability of the rate card
The marketing site’s /pricing path 404s and the real rate card lives on a separate cloud subdomain, so prospective buyers hit a dead end on the obvious URL. See bill shock and cost unpredictability.
2. Output-token predictability
With output rates 2–5x input and reasoning tokens billed as output, bills are hard to forecast. A token-estimator or per-request cost preview would reduce surprise charges for chain-of-thought workloads.
3. Opaque systems pricing
The entire RDU hardware business is sales-quoted with no indicative public number, which slows evaluation for buyers comparing against GPU-cloud alternatives that publish at least banded rates.
4. The card gate narrows the top of the funnel
As of August 14, 2026, evaluating SambaCloud requires adding a payment method before the first request — the earlier no-card $5 credit is gone. That’s a reasonable anti-abuse move, but it removes the lowest-friction way to try a hardware company’s API, likely costing some casual sign-ups who would have converted after a card-free test call. A short, rate-limited no-card sandbox (a handful of calls against a small model) would preserve most of the abuse protection while keeping a true zero-friction entry point.
Monetization stack & signals : how SambaNova builds & buys its revenue engine
Buys 2 Builds 1
Buys the self-serve money-movement layer (Stripe portal for cards/invoices, AWS Marketplace as the enterprise procurement rail) while its own per-token usage meter sits in front, feeding both. No sourced sign of a CRM/CPQ/rev-rec stack behind the sales-quoted RDU hardware — that quote-to-cash spine stays unconfirmed.
-
“Per-1M-token input/output rates metered against $5 of credits, then drawn down per token on the Developer tier — a usage meter feeding the Stripe payment layer.”
-
“The 'Manage Billing' link opens the Stripe customer portal; the Stripe billing portal shows 'No invoice history' — meaning none of these invoices are ever pushed to Stripe for payment.”
-
“SambaNova is available through the AWS Marketplace, enabling enterprises to streamline procurement and consolidate billing through their existing AWS account.”
Signals reviewed · derived from product docs
Key takeaways
- SambaNova is a hybrid model — public per-token usage pricing on the cloud API, sales-quoted contracts on RDU hardware. For the underlying model, see the introduction to usage-based pricing.
- Token rates span ~14x by model — from $0.22 input (gemma-4-31B-it, gpt-oss-120b) to $3.00 input (DeepSeek-V3.1/V3.2) — so model selection, not tier, drives the bill; MiniMax-M2.7 also now carries a $0.06/1M cached-input rate for repeated prompts.
- There’s a $0 free tier on inference — but as of August 2026 it requires adding a payment method to purchase pay-as-you-go credits; the earlier no-card $5 credit grant was removed.
- The differentiation is speed, not price — SambaNova runs open models on its own RDU silicon and sells fastest-inference at mid-pack token rates.
- The hardware business stays opaque — no public price for SambaStack/SambaManaged/DataScale; everything above the API is a sales conversation.
UBP implications
- A usage-based API can be a funnel for a non-usage product — even after the free tier stops being frictionless. SambaNova uses a metered inference API to generate demand for sales-quoted silicon, and kept that funnel intact on August 14, 2026 by card-gating the $0 Free plan rather than removing it outright — usage pricing as acquisition, not just monetization, even as the vendor tightens who gets to try it for free.
- The value metric need not be the cheapest unit. Pricing per token while competing on tokens-per-second shows a usage-based vendor can hold mid-pack unit prices if it differentiates on a quality dimension buyers can feel.
- Per-item pricing can replace tiered packaging. Letting each model set its own rate across a wide spread lets customers self-select cost vs. quality without bundles — a clean pattern for catalogs of fungible units.
Sources
- SambaNova Cloud pricing (per-token rate card) (accessed 2026-08-15)
- SambaNova Cloud plans (Free / Developer / Enterprise) (accessed 2026-08-15)
- SambaNova docs — SambaCloud supported models (independent confirmation of the current model lineup) (accessed 2026-08-15)
- SambaNova systems marketing site (systems sales-quoted; $11B raise announced July 8, 2026) (accessed 2026-07-23)
- SambaNova blog — SN50 RDU & Series E (accessed 2026-06-15)
Bottom line
SambaNova is a hybrid pricing story: a transparent, usage-based inference API (SambaNova Cloud) with a free tier and per-1M-token rates from $0.22 to $3.00 input (plus a new $0.06/1M cached-input rate on MiniMax-M2.7), bolted onto a sales-only RDU hardware business with no public price at all. The cloud rate card competes on speed rather than the cheapest token — SambaNova runs open models on its own silicon and sells fastest-inference — while the $0 Free tier funnels developers toward both pay-as-you-go usage and, eventually, enterprise systems deals — though since August 2026 that funnel starts with a payment method rather than a no-card credit grant. The things to watch are output-token costs and the opaque hardware pricing above the API. Browse the pricing blueprint for more fully-researched company profiles, or compare SambaNova against other Infrastructure, Compute & MLOps companies.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
No-card $5 free-credit grant removed from the Free plan
The SambaNova Cloud Free plan's 'no credit card required' $5 API credit grant is gone: the plans page now reads 'Add a payment method and purchase credits to run your first requests,' dropping the $5-credit, no-card-required, and 30-day-expiry language that was present as of August 11, 2026. Free is still $0 with no subscription fee, but there's no longer a way to call the API before adding a card. Token rates on the Pricing page are unchanged.
gemma-4-31B-it price cut reversed
August 2026 rate card: gemma-4-31B-it reverts from $0.22/$0.59 back to $0.38/$1.15 input/output, undoing the July 2026 cut after roughly three weeks. All five other models on the card (MiniMax-M2.7, DeepSeek-V3.1, DeepSeek-V3.2, gpt-oss-120b, Meta-Llama-3.3-70B) are unchanged; gpt-oss-120b is now the sole model at the $0.22 price floor.
Cached-input pricing added; gemma cut to $0.22; $11B raise
July 2026 rate card trimmed to six models: gemma-4-31B-it cut from $0.38/$1.15 to $0.22/$0.59; DeepSeek-V3.1-cb and DeepSeek-R1-Distill dropped; a new Cached Input Tokens column debuts with MiniMax-M2.7 at $0.06/1M cached. DeepSeek-V3.1/V3.2 stay $3.00/$4.50, Llama-3.3-70B $0.60/$1.20. SambaNova closed a $1B round at an $11B valuation (General Atlantic) on July 8, 2026.
Rate card spans $0.15 to $3.00 input across newer models
June 2026 SambaCloud rate card: DeepSeek-V3.1-cb $0.15/$0.75, gpt-oss-120b $0.22/$0.59, gemma-4-31B-it $0.38/$1.15, Meta-Llama-3.3-70B $0.60/$1.20, MiniMax-M2.7 $0.60/$2.40, DeepSeek-R1-Distill-Llama-70B $0.70/$1.40, DeepSeek-V3.1/V3.2 $3.00/$4.50. SN50 RDU and $350M Series E announced Feb 2026.
Sovereign-AI inference partnerships expand the footprint
SambaNova signs sovereign-AI inference deals (Argyll UK, Infercom Germany, OVHcloud EU, SouthernCrossAI Australia), extending the token-based cloud into region-specific clouds while keeping the published rate card.
SambaNova Cloud launches with a free developer tier
SambaNova opens a public, OpenAI-compatible inference API positioned on fastest-token-throughput for open models (Llama family), with a free tier and pay-as-you-go per-token billing rather than per-GPU-hour.
- · SambaNova was founded in 2017 by Stanford professors Kunle Olukotun and Christopher Ré with ex-Oracle exec Rodrigo Liang; its 2021 Series D ($676M, SoftBank-led) valued it above 5B.
- · Its pricing pitch isn't the cheapest token — it's the fastest. SambaNova runs open models on its own RDU silicon and routinely claims record tokens-per-second for Llama, DeepSeek, gpt-oss and Gemma.
- · The public rate card lives on cloud.sambanova.ai, not sambanova.ai/pricing — the marketing domain's /pricing path 404s, because the hardware business has no public price at all.
Questions & answers
- How does SambaNova's pricing work?
- SambaNova is hybrid. SambaNova Cloud (SambaCloud) is a public, usage-based inference API billed per 1M tokens, with a Free tier (as of August 2026, requires a payment method to purchase pay-as-you-go credits), a pay-as-you-go Developer tier, and a subscription-based Enterprise tier. Separately, SambaNova sells RDU-based hardware systems (SambaStack, SambaManaged) that are sales-quoted with no public rate card.
- How much does SambaNova Cloud cost per million tokens?
- As of August 2026, SambaCloud token rates range from $0.22 input / $0.59 output (gpt-oss-120b) to $3.00 input / $4.50 output for DeepSeek-V3.1 and V3.2. gemma-4-31B-it sits at $0.38 input / $1.15 output after a brief July cut to $0.22/$0.59 was reversed. Meta-Llama-3.3-70B-Instruct is $0.60 input / $1.20 output, and MiniMax-M2.7 is $0.60 / $2.40 with a $0.06/1M cached-input rate — the only model on the card to carry cached-input pricing.
- Does SambaNova have a free tier?
- As of August 2026, no longer in the no-card sense: SambaNova removed the $5 free-credit grant. The SambaNova Cloud Free plan still exists at $0, but now requires you to add a payment method and purchase pay-as-you-go credits before you can run requests; it includes access to Production models and community support. The hardware/systems business has no free tier and is sold through sales.
- Is SambaNova usage-based or subscription pricing?
- Both, depending on the product. The cloud inference API is pure usage-based (per-token, pay-as-you-go) on the Free and Developer tiers, shifting to subscription-based pricing on Enterprise for larger usage. The RDU hardware systems are sold as sales-quoted enterprise contracts.