Ask
All companies
technology

Baseten pricing

baseten.co facts checked analysis reviewed
Quick summary
Product
ML inference infrastructure — dedicated GPU deployments, Model APIs, and Truss framework
Industry
technology
Commits
Available (annual)
In this page
AI Summary
  • Baseten runs a pure-usage GPU-minute billing model for dedicated model deployments plus a separate per-token Model APIs catalog — both pay-as-you-go from the Basic tier with no monthly minimum.
  • Dedicated deployment per-minute rates as of 2026: T4 $0.01052, L4 $0.01414, A10G $0.02012, A100 80GB $0.06667, H100 MIG 40GB $0.0625, H100 80GB $0.10833, B200 180GB $0.16633 — all billed only for active inference time, with scale-to-zero idle replicas free.
  • Model APIs charge per million tokens across a fourteen-model catalog: GPT OSS 120B at $0.10/$0.50, DeepSeek-V4-Flash-0731 at $0.13/$0.26, GLM-5.3-Flash at $0.15/$0.50, Inkling-Small at $0.50/$1.20, GLM 4.7 at $0.60/$2.20, Kimi K2.6 at $0.95/$4.00, Inkling at $1.00/$4.05, DeepSeek V4 Pro 0813 at $1.32/$3.96, GLM-5.2 at $1.40/$4.40, DeepSeek V4 Pro at $1.74/$3.48, GLM-5.2 Fast at $2.10/$6.60, and flagship Kimi K3 at $3.00/$15.00 — with a published cache-input column roughly 80–90% below the standard input rate on most models. The catalog grew from ten to twelve models on 2026-08-04 with the addition of DeepSeek-V4-Flash-0731 and Inkling-Small (the existing DeepSeek V4 SKU renamed DeepSeek V4 Pro at unchanged rates), then to fourteen models on 2026-08-28 with GLM-5.3-Flash and a second dated DeepSeek V4 Pro 0813 variant — no rate on any of the twelve prior SKUs moved in that update.
  • New accounts receive free credits; the Basic tier carries no monthly fee and is SOC 2 Type II + HIPAA-compliant out of the box.
  • Pro and Enterprise tiers add priority GPU access, dedicated compute reservations, higher rate limits, custom SLAs, BYOC self-hosted deployment, data-residency controls, and Slack/Zoom support — priced via annual usage commitments rather than published tier rates.
  • Baseten raised a $75M Series C in February 2024 led by IVP at an ~$825M post-money valuation; a Series D round was reported in late 2025 at >$2B.
Pricing summary
Baseten 2026 — Pay-per-minute GPU + per-token Model APIs
Basic ($0/mo, pay as you go) → Pro (volume discounts) → Enterprise (custom SLAs, self-host / hybrid)
Basic
$0 /mo + usage
Developers, startups, prototypes
Get a quote
Enterprise
Custom
Regulated industries, mission-critical inference
GPU per-minute
From $0.01052 /min (T4)
Dedicated single-tenant deployments
Model APIs
From $0.10 /1M tokens
Per-token multi-tenant endpoints
No monthly fee on Basic. Dedicated GPU billed per minute (active inference only); idle replicas in scale-to-zero state are free. Volume discounts on Pro and Enterprise require quote-based commitments.

About

Baseten is a San Francisco-based ML infrastructure company founded in 2019 by Tuhin Srivastava, Amir Haghighat, and Philip Howes — all formerly at Gumroad. The product is a managed inference platform that lets ML and product teams deploy custom or open-source models behind production-grade endpoints without operating a GPU fleet, Kubernetes cluster, or autoscaler. The interface is opinionated: model code is packaged with Truss (Baseten’s open-source framework), deployed to dedicated single-tenant GPU instances, and exposed via REST or gRPC endpoints with automatic scale-to-zero, configurable warm-pool windows, and per-replica observability.

By 2026 Baseten serves Writer, Descript, Patreon, Robust Intelligence, Picnic Health, and roughly a thousand other paying customers spanning enterprise compliance-sensitive workloads (HIPAA-regulated healthcare AI, financial-services NLP) and high-growth AI-native startups serving sub-second inference SLAs. The company raised a $75M Series C in February 2024 led by IVP at an ~$825M post-money valuation, with a later Series D round at >$2B reported in late 2025; the site’s top banner rotates between funding and product news — an “Announcing our Series F” banner in July 2026 gave way to a “Try the new DeepSeek V4 Flash today” promo by early August 2026, then to “Try the new GLM-5.3 Flash today” by late August 2026. Logos displayed above the pricing tiers include Abridge, Ambience, Canopy Labs, Clay, ClickUp, Cursor, Decagon, and Descript.

Baseten competes with hyperscaler inference platforms (AWS Bedrock, Google Vertex AI, Azure ML), specialized inference clouds (Fireworks AI, Together AI, Replicate), and serverless GPU providers (Modal, RunPod). Its differentiation is the combination of per-minute transparent GPU billing, model-agnostic deployment (any PyTorch or TensorFlow model, not just a predefined catalog), and a Truss-based developer experience that handles cold-start optimization, request batching, and concurrency tuning without per-model engineering effort.


Pricing summary : How Baseten’s per-minute GPU + per-token APIs stack works

Baseten runs two parallel pricing surfaces that share the same platform credits balance. Dedicated deployments are billed by GPU-minute on whichever instance type the customer selects (T4 through B200), with no markup for the platform layer beyond the published rate — the price is the same whether the model is open-source, a customer’s proprietary checkpoint, or a fine-tuned variant. The rate card carries a Minute/Hour toggle so the same rate is shown at either granularity, and a separate on-demand Training rate card and CPU-instance rate card sit alongside the GPU tables. Model APIs are multi-tenant per-token endpoints for a catalog that grew from twelve to fourteen models on 2026-08-28 with the addition of GLM-5.3-Flash and DeepSeek V4 Pro 0813 (no rate changed on any of the twelve prior SKUs) — the full roster is Kimi K3, Kimi K2.6, Kimi K2.7 Code, Inkling-Small, Inkling, GLM-5.2, GLM-5.2 Fast, GLM-5.3-Flash, GLM 4.7, NVIDIA Nemotron 3 Ultra, DeepSeek-V4-Flash-0731, DeepSeek V4 Pro, DeepSeek V4 Pro 0813, and GPT OSS 120B — priced per million input/output tokens with a published Cache Input column that runs roughly 80–90% below the standard input rate.

The Basic tier is labelled “$0 per month, pay as you go” — customers pay only for active usage. Pro and Enterprise tiers do not change the per-minute or per-token rates publicly; both show only “Volume discounts available” behind a “Get a quote” CTA, and they unlock priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, custom SLAs, and self-host / hybrid deployment. This dual-track pricing — transparent self-serve consumption plus quote-based commitments — mirrors the pure-usage + commitment hybrid increasingly common across AI infrastructure companies.

What makes this different: Baseten publishes per-minute pricing rather than per-hour, which makes scale-to-zero economics legible. A model that bursts for 4 minutes per request on a $0.10833/min H100 is billed for exactly those 4 minutes — the kind of granular cost-modeling that AWS Bedrock and Vertex AI deliberately hide behind aggregate per-month bills.


Pricing by product

Dedicated GPU deployments (per-minute, single-tenant)

Baseten’s pricing page ships a Minute / Hour toggle on the dedicated-deployment rate card, so both columns below are published rates read directly from the page — not derived. Billing is always metered to the minute; the per-hour figures are the same rate viewed at hourly granularity.

InstanceVRAMPer-minute ratePer-hour rateBest for
T416 GiB$0.01052$0.6312Low-cost embeddings, small model inference
L424 GiB$0.01414$0.8484Mid-range models, video inference
A10G24 GiB$0.02012$1.20727B–13B model inference, image gen
A100 80GB80 GiB$0.06667$4.0030B–70B model inference, training
H100 MIG 40GB40 GiB$0.0625$3.75Cost-optimized H100 partition
H100 80GB80 GiB$0.10833$6.50Frontier model inference, low-latency serving
B200 180GB180 GiB$0.16633$9.98Largest open models, multi-model serving

CPU-only instances are also billed per minute for pre/post-processing and CPU model serving, and carry the same Minute / Hour toggle:

CPU instanceSpecPer-minute ratePer-hour rate
1x21 vCPU, 2 GiB RAM$0.00058$0.0348
1x41 vCPU, 4 GiB RAM$0.00086$0.0516
2x82 vCPUs, 8 GiB RAM$0.00173$0.1038
4x164 vCPUs, 16 GiB RAM$0.00346$0.2076
8x328 vCPUs, 32 GiB RAM$0.00691$0.4146
16x6416 vCPUs, 64 GiB RAM$0.01382$0.8292

Both rate cards carry a “Volume discounts available” badge, and each table is footnoted “Talk to sales about compute in other countries and regions” — capacity outside the default footprint is a sales conversation, not a self-serve selection.

Training (on-demand, per-minute)

A separate Training rate card (“on-demand compute, devex, and infrastructure for your training jobs”) is included in Basic and priced at the same per-minute GPU rates as dedicated inference: T4 $0.01052, L4 $0.01414, A10G $0.02012, A100 80 GiB $0.06667, H100 MIG 40 GiB $0.0625, H100 80 GiB $0.10833, B200 180 GiB $0.16633. It carries its own Minute / Hour toggle and the same “Volume discounts available” badge; no CPU-instance row is offered for training.

Model APIs (per-token, multi-tenant)

Prices are per 1M tokens. The pricing page publishes a dedicated Cache Input column — cached prompt-prefix tokens are billed at a large discount to the standard input rate (roughly 80–90% off, e.g. DeepSeek V4 Pro input drops from $1.74 to $0.145). The catalog stood at twelve models after the 2026-08-04 rename of DeepSeek V4 to DeepSeek V4 Pro; by 2026-08-28 it had grown to fourteen with the addition of GLM-5.3-Flash ($0.15 input / $0.03 cache / $0.50 output — flagged by a homepage “Try the new GLM-5.3 Flash today” banner) and DeepSeek V4 Pro 0813 ($1.32 / $0.132 / $3.96, a dated variant of DeepSeek V4 Pro with a lower input rate but higher output rate). No rate on any of the twelve prior SKUs changed.

ModelInput ($/1M)Cache input ($/1M)Output ($/1M)
Kimi K3$3.00$0.30$15.00
GLM-5.2 Fast$2.10$0.21$6.60
DeepSeek V4 Pro$1.74$0.145$3.48
GLM-5.2$1.40$0.14$4.40
DeepSeek V4 Pro 0813$1.32$0.132$3.96
Inkling$1.00$0.17$4.05
Kimi K2.7 Code$0.95$0.16$4.00
Kimi K2.6$0.95$0.16$4.00
GLM 4.7$0.60$0.12$2.20
NVIDIA Nemotron 3 Ultra$0.60$0.12$2.40
Inkling-Small$0.50$0.10$1.20
GLM-5.3-Flash$0.15$0.03$0.50
DeepSeek-V4-Flash-0731$0.13$0.028$0.26
GPT OSS 120B$0.10— (no cache rate published)$0.50

Tier features (non-usage)

TierMonthly feeVolume discountDeployment optionsIncluded / addedSupport
Basic$0 per month, pay as you goNot offeredBasetenDedicated deployments, Model APIs, Training, fast cold starts, SOC 2 Type II and HIPAA compliantEmail and in-app chat
ProQuote”Volume discounts available”BasetenEverything in Basic plus priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, hands-on engineering expertiseDedicated support on Slack and Zoom
EnterpriseQuote”Volume discounts available”Baseten, your VPC, hybridEverything in Pro plus custom SLAs, self-host deployments, on-demand flex compute, use existing cloud commitments, full control over data residency, advanced security and compliance, custom global regions, advanced RBAC with TeamsDedicated support on Slack and Zoom

Sales motions across products: PLG / self-serve for Basic (per-minute GPU/CPU + per-token Model APIs + on-demand Training), sales-led for Pro and Enterprise commits and BYOC.


Hidden costs : What Baseten customers actually pay beyond the base GPU rate

Archetype A: AI-native startup running one 7B model on A10G with bursty traffic

A growth-stage startup serving ~50,000 requests/day, average 800ms inference time, with traffic concentrated in business hours:

Line itemMonthly cost
A10G compute (4h/day active × 30 × $0.02012 × 60)$145
Warm-pool retention (avoid cold starts business hours, ~6h/day)$72
Outbound bandwidth (negligible at this scale)<$5
Estimated total~$220/month

The warm-pool premium roughly doubles compute cost — but eliminates 3–8 second cold starts that would push p95 latency above customer SLAs. Most production deployments end up paying a 30–60% warm-pool premium over pure scale-to-zero.

Archetype B: Mid-market team running mixed Model APIs + dedicated H100

A team using Model APIs for general-purpose Q&A and a dedicated H100 for a fine-tuned model:

Line itemMonthly cost
DeepSeek V4 Model API (50M input at $1.74 + 15M output at $3.48)$139
Cached input savings (40% of input cached at $0.145 instead of $1.74)-$32
Dedicated H100 80GB (8h/day × 30 × $0.10833 × 60)$1,560
Mission Critical SLA add-on (Enterprise)Quote
Estimated total~$1,670/month + SLA quote

H100 dedicated compute dominates the bill — and the customer is paying for warm-pool retention to maintain low latency. The cached input discount is steep (DeepSeek V4’s $0.145 cache rate is ~92% below its $1.74 input rate) but only bites on workloads with repetitive prompt prefixes (RAG, agent loops); ad-hoc query workloads see negligible cache hits. The line worth stress-testing is the model row itself: after the 2026-07-21 catalog sweep retired four SKUs, a Model APIs forecast built on the model a team originally integrated can silently describe a model that is no longer on the rate card — re-run it against the ten currently published SKUs.

Want to estimate your own Baseten bill? Use the Baseten pricing calculator to model dedicated GPU cost by instance and active minutes, plus Model API spend by token volume.


Pricing evolution : Baseten’s pricing history from no-code serving to per-minute GPU transparency

Cadence

QuarterPrice changesProduct / SKU additionsNotes
2019 Q101Company founded; initial no-code model-serving product
2022 Q211Series B; Truss open-sourced; pricing shifted to per-minute GPU
2023 Q401Model Library launched — one-click open-source model deployments
2024 Q100Series C ($75M IVP-led); no public price changes
2024 Q301Model APIs (multi-tenant per-token) launched
2024 Q410H100 80GB published at $0.10833/min as a transparency play
2025 Q101B200 instances added at $0.16633/min
2025 Q301Self-host (BYOC) deployment option launched
2026 Q110Cached input pricing added to Model APIs
2026 Q3172026-07-21 — Model APIs catalog pruned 11 SKUs → 8 (GLM 5.1, GLM 5, Kimi K2.5, NVIDIA Nemotron 3 Super retired); Inkling added at $1.00 in / $0.17 cache / $4.05 out; 2026-07-29 — catalog grew to 10 with flagship Kimi K3 ($3.00/$0.30/$15.00) and GLM-5.2 Fast ($2.10/$0.21/$6.60) added, and GLM-5.2’s cache-input rate cut 46% to $0.14; 2026-08-04 — catalog grew to 12 with DeepSeek-V4-Flash-0731 ($0.13/$0.028/$0.26) and Inkling-Small ($0.50/$0.10/$1.20) added, and the existing DeepSeek V4 SKU renamed DeepSeek V4 Pro at unchanged rates; 2026-08-28 — catalog grew to 14 with GLM-5.3-Flash ($0.15/$0.03/$0.50) and DeepSeek V4 Pro 0813 ($1.32/$0.132/$3.96) added, no existing SKU repriced; all GPU, CPU, Training and tier surfaces unchanged across all four events

Tracked range: 2019 Q1–2026 Q3. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).

Notable changes

  • 2022-04-19 — Truss open-sourced; pricing pivoted from seat-based no-code SaaS to per-minute GPU consumption — the defining structural choice that still shapes Baseten’s positioning.
  • 2023-11-10 — Model Library launched with pre-deployed open-source models (Stable Diffusion, Whisper, Llama 2, Mistral); per-minute rate of underlying instance, no model markup.
  • 2024-08-05 — Model APIs launched, adding pure pay-per-token multi-tenant endpoints to complement per-minute dedicated.
  • 2024-11-18 — H100 published at $0.10833/min — Baseten leaned into transparency as differentiation against opaque hyperscaler markup.
  • 2025-09-22 — Self-host / BYOC option launched for enterprises with strict data-residency or compute-cost requirements.
  • 2026-02-14 — Cached input pricing added to Model APIs; brought parity with first-party caching offerings from OpenAI and Anthropic.
  • 2026-07-21 — Model APIs catalog cut from eleven models to eight in a single sweep: GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super retired, Inkling added at $1.00 input / $0.17 cache / $4.05 output. No rate on any surviving SKU moved, and the dedicated GPU, CPU, Training and Basic/Pro/Enterprise surfaces were all unchanged — this was catalog pruning, not repricing.
  • 2026-07-29 — Eight days after that pruning, the Model APIs catalog grew back to ten models: flagship Kimi K3 launched at $3.00 input / $0.30 cache / $15.00 output (72% above the prior input ceiling and over 3x the prior output ceiling), and GLM-5.2 Fast joined at $2.10 / $0.21 / $6.60. The same capture cut GLM-5.2’s own cache-input rate 46%, from $0.26 to $0.14, while its $1.40 input and $4.40 output rates held steady. GPU/CPU/Training rates and the Basic/Pro/Enterprise tier structure were unchanged.
  • 2026-08-04 — Six days after that, the Model APIs catalog grew from ten to twelve models: DeepSeek-V4-Flash-0731 launched at $0.13 input / $0.028 cache / $0.26 output — the platform’s cheapest model to date, undercutting even the pre-pruning cache-rate floor — alongside Inkling-Small at $0.50 / $0.10 / $1.20. The existing DeepSeek V4 SKU was renamed DeepSeek V4 Pro in the same capture, with its $1.74 / $0.145 / $3.48 rates unchanged. The homepage banner switched from an “Announcing our Series F” funding promo to a “Try the new DeepSeek V4 Flash today” product promo — the first time the site’s top banner led with a Model APIs launch rather than a funding milestone.
  • 2026-08-28 — Three weeks after the DeepSeek V4 Pro rename, the Model APIs catalog grew from twelve to fourteen models: GLM-5.3-Flash launched at $0.15 input / $0.03 cache / $0.50 output, and DeepSeek V4 Pro 0813 joined as a second, dated DeepSeek V4 variant at $1.32 / $0.132 / $3.96 — a 24% lower input rate than DeepSeek V4 Pro’s $1.74 but a 14% higher output rate than its $3.48. No rate on any of the twelve prior SKUs changed. The homepage banner switched again, from the DeepSeek V4 Flash promo to a “Try the new GLM-5.3 Flash today” promo — the second consecutive banner cycle built around a Model APIs launch rather than a funding milestone.

July–August 2026’s Model APIs churn, in detail

Three of the four retired models sat at the cheap end of the rate card, so the practical effect is a floor raise rather than a price rise. With NVIDIA Nemotron 3 Super ($0.30 in / $0.06 cache / $0.75 out) and Kimi K2.5 ($0.60 / $0.12 / $3.00) gone, the lowest published input rate outside GPT OSS 120B ($0.10, and the only SKU with no cache rate at all) is now $0.60, and the cheapest published cache-input rate doubled from $0.06 to $0.12. A team that had standardised its cheap-tier RAG or classification traffic on Nemotron 3 Super now has no like-for-like replacement on the rate card: the nearest survivor, Nemotron 3 Ultra, costs 2× the input rate and 3.2× the output rate.

Inkling ($1.00 / $0.17 / $4.05) landed at what was then the premium end, roughly between Kimi K2.6 and GLM 5.2 — so the 07-21 swap traded four SKUs of price spread for one more frontier-class option, briefly narrowing the catalog’s input-rate spread to $0.10–$1.74. Baseten runs a documented deprecation surface in its Model APIs docs, which is what makes a four-SKU sweep read as scheduled housekeeping — but the pricing page itself gave no advance notice, so the removal was visible to a buyer only after the fact.

Eight days later, on 2026-07-29, the catalog moved again — this time upward and outward rather than by subtraction. Kimi K3 reset the ceiling entirely: at $3.00 input / $15.00 output it is 72% pricier on input and more than 3x pricier on output than the previous top SKU (DeepSeek V4), stretching the catalog’s input-rate spread from $0.10–$1.74 back out to $0.10–$3.00 in a single capture. GLM-5.2 Fast slotted in just below it as a second premium option. In the same breath, Baseten cut an existing model’s economics — GLM-5.2’s cache-input rate fell 46% to $0.14 — with no corresponding change to its input or output rate and no new SKU required to justify it. Read together, the two events show that Baseten’s Model APIs rate card is not settling into stability after the July pruning; it is actively and continuously managed in both directions — retiring cheap SKUs one week, adding an expensive flagship and quietly re-pricing a survivor’s cache rate the next. A buyer budgeting against this catalog should treat the rate card itself, not any single price point, as the thing to monitor.

A third data point landed on 2026-08-04, just six days later, and it looks different from the two before it. Unlike 07-21 (net contraction) and 07-29 (pure addition), this event adds two SKUs and renames a third. DeepSeek-V4-Flash-0731 joined at $0.13 input / $0.028 cache / $0.26 output — comfortably the cheapest model Baseten has ever listed, and the new cache-rate floor: the catalog’s cheapest published cache-input rate, which had doubled from $0.06 to $0.12 in the 07-21 pruning, now drops to $0.028 — more than 50% below even the pre-pruning floor. Inkling-Small joined alongside it at $0.50 / $0.10 / $1.20, a smaller sibling to the existing Inkling line. More notable than either addition: the original DeepSeek V4 entry was simultaneously renamed DeepSeek V4 Pro, at unchanged rates. It is the first time in this run of changes that Baseten has altered a model’s identifier on the rate card without touching its price or retiring it — a rename, not a repricing or a retirement. The catalog is now twelve models, net growth of four since the 07-21 low of eight, with zero further retirements since. The homepage banner also flipped for the first time from a funding story (“Announcing our Series F”) to a product story (the DeepSeek V4 Flash promo) — a signal that Baseten is now treating cheap-tier Model API additions as a marketable growth lever, not just quiet catalog housekeeping.

A fourth data point landed on 2026-08-28, three weeks after the DeepSeek V4 Pro rename, and it is the simplest of the four: pure addition, with no retirement, repricing, or rename attached. GLM-5.3-Flash joined at $0.15 input / $0.03 cache / $0.50 output — its $0.03 cache rate sits just above the $0.028 floor DeepSeek-V4-Flash-0731 set on 2026-08-04, so the catalog’s cheapest cache rate held rather than moved again this time. DeepSeek V4 Pro 0813 joined alongside it as a second, dated DeepSeek V4 variant: its $1.32 input rate undercuts the original DeepSeek V4 Pro’s $1.74 by 24%, but its $3.96 output rate runs 14% above DeepSeek V4 Pro’s $3.48 — the first Model APIs addition on this rate card to cut one rate while raising another within the same model family. The homepage banner moved again too, from the DeepSeek V4 Flash promo to a GLM-5.3-Flash promo — the second straight cycle where Baseten marketed a cheap-tier Model APIs launch on its homepage rather than a funding milestone. Four catalog events in five weeks (07-21, 07-29, 08-04, 08-28) is no longer a pruning-then-settling story; it’s evidence the rate card is a permanently live surface a buyer has to re-check on a cadence measured in weeks, not quarters.


What’s unique : Baseten’s distinctive pricing mechanics

1. Per-minute (not per-hour) billing makes scale-to-zero economics visible. Most cloud GPU offerings — AWS Sagemaker, Vertex AI, Azure ML — bill per hour with billing-minute rounding. Baseten’s per-minute granularity means a 90-second burst on an H100 costs $0.16, not a rounded-up $6.50/hour. This makes usage forecasting materially more accurate for bursty workloads and exposes the real cost of warm-pool retention versus pure scale-to-zero.

2. No model markup on dedicated deployments — only the GPU rate. When a customer runs Llama 3.3 70B on an A100 via Baseten’s Model Library, they pay $0.06667/min — the same A100 rate as if they were running their own proprietary model. Most managed inference platforms charge a per-token markup on hosted open-source models on top of the underlying compute cost. Baseten’s flat per-minute rate makes the platform layer free of tier-based pricing distortion.

3. Dual-SKU split: per-minute dedicated AND per-token multi-tenant in the same product. Baseten lets customers use Model APIs (per-token, multi-tenant) for low-volume general-purpose queries and dedicated deployments (per-minute, single-tenant) for sustained high-QPS or proprietary models — in the same workspace, with the same billing balance. Most competitors force a choice between platforms (Replicate vs Fireworks, for example). This hybrid usage model reduces vendor sprawl for AI-native teams. The division of labour also explains why the multi-tenant catalog is deliberately short and actively managed rather than static — pruned from eleven models to eight on 2026-07-21, grown back to ten on 2026-07-29 with flagship Kimi K3 and GLM-5.2 Fast, then to twelve on 2026-08-04 with DeepSeek-V4-Flash-0731 and Inkling-Small (the existing DeepSeek V4 SKU renamed DeepSeek V4 Pro in the same capture), and to fourteen on 2026-08-28 with GLM-5.3-Flash and a second dated variant, DeepSeek V4 Pro 0813. Model APIs exist to cover convenience traffic on a curated, moving set of frontier open-weight models; anything Baseten will not keep warm multi-tenant is meant to move to a dedicated per-minute deployment, where the customer rather than the catalog decides the model’s lifespan.

4. Cached input pricing applies to multi-tenant Model APIs — but the discount travels with the SKU, not the account. Cached-input discounts are normally exclusive to first-party providers (OpenAI, Anthropic) where the cache is implementation-controlled. Baseten ships a published Cache Input column on hosted DeepSeek, Kimi, GLM and Nemotron endpoints at 80–90% below the standard input rate, so customers using Baseten’s Model APIs as a DeepSeek proxy get caching parity with first-party DeepSeek without negotiating a separate contract. The catch surfaced on 2026-07-21: because the discount is a per-model column rather than an account-level entitlement, retiring a cheap SKU also retires its cheap cache rate — pruning Nemotron 3 Super and Kimi K2.5 doubled the catalog’s cheapest cache-input rate from $0.06 to $0.12 without any published rate changing. The same lever cuts the other way too: on 2026-07-29 Baseten cut GLM-5.2’s own cache-input rate 46% (from $0.26 to $0.14) with zero change to its input or output rate — a per-model discount that can improve unilaterally, at Baseten’s discretion, just as easily as it can vanish. On 2026-08-04 the floor moved again, further than either prior event: DeepSeek-V4-Flash-0731 launched with a $0.028 cache-input rate — not just a recovery from the post-pruning $0.12 floor but a new low more than 50% below the original pre-pruning floor of $0.06. A metric this volatile (doubling, then dropping to less than half its starting point, inside six weeks) is not something a customer can budget against once and forget; it has to be re-checked at the SKU level every time the catalog changes. The 2026-08-28 additions left that floor untouched — GLM-5.3-Flash’s $0.03 cache rate and DeepSeek V4 Pro 0813’s $0.132 both sit above it — but that’s a data point about this one event, not a new floor to assume holds going forward.

5. Self-host (BYOC) as an enterprise unlock without abandoning the platform. Baseten’s BYOC option lets large enterprises retain Truss-based deployment, scale-to-zero, observability, and Baseten’s control plane — but use their own AWS, GCP, or Azure GPU reservations and their own VPC for data residency. This platform-license + customer-compute model is unusual in inference middleware and addresses both compliance and committed-spend savings simultaneously.


Strengths & weaknesses

StrengthsWeaknesses
Transparent per-minute GPU pricing across all instance types (T4–B200)Pro and Enterprise tier prices are not published — must contact sales for discounts
No model markup on dedicated deployments — the GPU rate is the pricePer-minute H100 rate ($6.50/hr equiv) is 1.5–2× raw AWS on-demand H100 list
Scale-to-zero with configurable warm-pool eliminates idle billingCold-start latency on scale-to-zero (3–8s) requires warm-pool tuning for latency SLAs
Truss framework supports any PyTorch or TensorFlow model, not a fixed catalog — dedicated deployments are immune to catalog churnMixing Model APIs and dedicated deployments requires manual cost-modeling, and the Model APIs catalog itself churns (11 SKUs → 8 on 2026-07-21, 8 → 10 on 2026-07-29, 10 → 12 on 2026-08-04 including a SKU rename, 12 → 14 on 2026-08-28) with no advance notice or changelog on the pricing page
SOC 2 Type II + HIPAA out of the box on Basic tierMission Critical SLA (99.95%) is Enterprise-only — no published SLA on Basic
BYOC option preserves Baseten platform while using customer-owned GPU reservationsGeographic regions limited compared to hyperscalers; APAC residency requires Enterprise + BYOC

Billing UX : Baseten’s account controls and payment experience

  • “Billing and usage” console section — Baseten’s docs carry a dedicated Billing and usage area under Account, alongside Rate limits and budgets for Model APIs — the two surfaces where consumption is inspected and capped.
  • Rate limits and budgets (Model APIs) — A named docs control for setting per-workspace Model API rate limits and spend budgets, distinct from the dedicated-deployment path which has no equivalent cap.
  • Billing webhooks — A first-class Billing webhooks integration in the docs pushes billing events out of Baseten, matching the Frontier Gateway design of sending usage data out-of-band to an external billing provider.
  • Autoscaling controls that drive the bill — Per-deployment configuration of minimum replicas, maximum replicas, concurrency targets, and scale-down delay. The docs state models “scale to zero when idle, eliminating costs during quiet periods, and scale up within seconds when traffic arrives” — the scale-down delay is the direct dial between cold-start latency and idle spend.
  • Minute / Hour rate-card toggle — The pricing page itself lets buyers flip both the dedicated-deployment and Training rate cards between per-minute and per-hour display, so the same rate can be modelled at either granularity before signing up.
  • “Get a quote” gate on Pro and Enterprise — Neither paid tier shows a number; both display only “Volume discounts available” with a quote CTA, so discount schedules never appear self-serve.
  • Regional capacity via sales — Both compute rate cards are footnoted “Talk to sales about compute in other countries and regions”, making non-default regions a contract conversation rather than a console toggle.
  • Advanced RBAC with Teams — Listed as an Enterprise-only entitlement on the pricing page; workspace role separation is not part of Basic or Pro.
  • Frontier Gateway for downstream billing — For AI labs reselling their own hosted model, the docs position Frontier Gateway as the managed gateway with “federated keys and per-customer billing” — Baseten’s metering surfaced as the customer’s own billing input.

Strategic wins : Why Baseten’s pricing decisions worked

1. Per-minute transparency as the GTM wedge against hyperscaler opacity

Baseten’s decision to publish per-minute H100 and B200 rates created a sharp competitive contrast against AWS Bedrock, Vertex AI, and Azure ML — all of which embed inference costs inside aggregate monthly bills with no per-request visibility. For AI engineering leaders forecasting unit economics, Baseten’s transparent rate card lets them model cost-per-inference and amortize against revenue-per-query before signing a contract. This transparency-as-positioning is unusual in B2B infrastructure and likely accelerated mid-market adoption.

2. Truss as an open-source moat for dedicated-deployment pricing

By open-sourcing Truss in 2022, Baseten created a free-to-use packaging framework that became the de-facto standard for model deployment among ML teams. Customers who built on Truss for local development naturally migrated to Baseten for production hosting — making per-minute dedicated pricing the path of least resistance. This PLG via developer tooling drove growth without paid acquisition spend through 2023–2024.

3. Dual SKU (per-minute + per-token) caught both ends of the workload spectrum

Most inference platforms force a choice: pay per-token for multi-tenant (Replicate, Fireworks) or pay per-hour for dedicated (Modal, RunPod). Baseten’s combined offering — per-minute dedicated AND per-token Model APIs in one workspace — captures both bursty low-volume queries (cheap on Model APIs) and sustained high-QPS production (cheap on dedicated). The multi-dimensional usage model maximizes wallet share within each customer. The Model APIs side of that split keeps expanding at the cheap end too — the catalog grew from ten to twelve models on 2026-08-04 with DeepSeek-V4-Flash-0731 and Inkling-Small, both undercutting every model on the card except GPT OSS 120B’s flat $0.10 input rate, which pulls more low-volume convenience traffic onto the metered per-token SKU rather than pushing it toward a dedicated deployment. The catalog kept growing three weeks later, too: 2026-08-28 added GLM-5.3-Flash — the third-cheapest input rate on the card, behind only GPT OSS 120B and DeepSeek-V4-Flash-0731 — and DeepSeek V4 Pro 0813, a second DeepSeek V4 variant trading a lower input rate for a higher output rate, giving buyers a genuine within-family choice rather than a single fixed price point.

4. BYOC as the enterprise upsell unlock without abandoning the platform

Baseten’s 2025 launch of self-hosted (BYOC) deployment let large enterprises keep Truss, scale-to-zero, and observability while using their own GPU reservations and VPC. This addresses the two largest enterprise blockers — data residency and committed-spend optimization — without forcing the customer onto a different product. It is a textbook enterprise pricing unlock that captures large customers who would otherwise self-build on Kubernetes.


Areas to improve : Gaps in Baseten’s pricing approach

1. Pro tier pricing should be published

Today the per-minute and per-token rates are identical on Basic, Pro, and Enterprise — the tier difference is volume discount magnitude and support level. By hiding the Pro discount schedule behind a sales call, Baseten loses self-serve growth-stage customers who want to model expansion economics without a sales conversation. Publishing a tiered volume-discount schedule (e.g., 10% off above $5K/mo, 20% off above $25K/mo) would let mid-market customers self-qualify into Pro without sales friction.

2. Cold-start cost is invisible until you hit it

Scale-to-zero is positioned as a cost-saver — and it is — but the latency cost of cold starts (3–8 seconds on a large model) requires customers to either tune warm-pool retention manually or accept latency SLA misses. The pricing page does not show warm-pool premium calculations or cold-start probability curves; customers discover this only after deploying. Adding a cold-start economics calculator to the pricing page would set expectations and likely increase Pro tier conversions for latency-sensitive workloads.

3. No unified spend forecast across Model APIs + dedicated

Customers using both Model APIs and dedicated deployments must manually combine per-token and per-minute forecasts. Baseten’s console shows historical spend by SKU but does not project future spend across mixed workload types. A unified usage forecasting view — token volume + active GPU minutes projected forward — would materially reduce bill-shock anxiety for cost-sensitive AI engineering teams.

4. Pricing page lacks egress and bandwidth detail

For high-volume inference workloads serving large image, audio, or video payloads, network egress can become a meaningful cost line. Baseten’s pricing page does not break out bandwidth pricing — customers learn the rate from invoices. Making egress pricing explicit (and ideally bundling a generous free egress allowance) would reduce a recurring source of surprise bills.

5. Model API retirements should be dated on the rate card, not just in the docs

On 2026-07-21 four SKUs — GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super — disappeared from the Model APIs table between one capture and the next, with the cheapest cache-input rate doubling from $0.06 to $0.12 as a side effect. Baseten does maintain a Deprecation page in its Model APIs docs, but the pricing page a buyer actually budgets against carries no sunset dates and no “retiring on” markers, so a team can standardise on a cheap SKU one fortnight and find it gone the next. The fix is cheap and entirely presentational: mark end-of-life models on the rate card with a retirement date, and commit publicly to a minimum notice window (30–60 days) plus a named migration target for each retiring SKU. That converts an unbounded bill-shock and re-platforming risk into a scheduled, plannable one — and costs Baseten nothing, since the deprecation policy already exists.

The churn isn’t one-directional, which widens the ask. Eight days after that sweep, on 2026-07-29, Baseten added Kimi K3 and GLM-5.2 Fast and cut GLM-5.2’s cache-input rate 46% — additions and a favorable repricing, not a retirement, but still invisible on the rate card itself until a buyer happens to re-check it. A lightweight “What changed” changelog anchored directly to the pricing page — covering additions and rate adjustments as well as retirements — would let buyers track the catalog as a whole instead of re-diffing the table after the fact.

The 2026-08-04 update adds a fourth failure mode a changelog would catch: renames. Baseten renamed the DeepSeek V4 SKU to DeepSeek V4 Pro in the same capture that added DeepSeek-V4-Flash-0731 — a sensible disambiguation once two DeepSeek V4 variants exist side by side, but a silent one. Nothing on the pricing page flags that “DeepSeek V4” and “DeepSeek V4 Pro” are the same model at the same price; a customer who hardcoded the old name into a config file, internal dashboard, or support ticket has no on-page signal that the rename happened, and could plausibly mistake DeepSeek V4 Pro for a new, untested SKU rather than their existing model under a new label. The fix extends the changelog proposal above: log renames alongside additions, retirements, and rate adjustments, and keep the old name as a searchable alias for at least one release cycle.

A fourth event, on 2026-08-28, was purely additive — GLM-5.3-Flash and DeepSeek V4 Pro 0813 joined with no retirement, repricing, or rename attached — but it landed with the same silence as the other three: no changelog entry, no “just added” marker, nothing beyond a changed row count on the rate card. Four unannounced catalog events in five weeks is the strongest evidence yet that this fix belongs on the near-term roadmap, not after the next retirement catches a customer off guard.


Monetization stack & signals : how Baseten builds & buys its revenue engine

Buys 3 Builds 1 9 open roles

The read — where the monetization investment is going

Baseten builds the meter but buys the billing layer: its in-house Frontier Gateway tracks usage per API key, then ships it out-of-band to an Orb-backed billing system. A growing RevOps/deal-desk org on Salesforce is wiring up the sales-led enterprise commit motion.

Stack — build vs buy
Builds in-house · 1
  • Frontier Gateway (in-house metering & rate-limiting) In-house build Blog May 2026

    “Token or character consumption is tracked per API key. Usage data is sent out-of-band to allow you to plug directly into your billing provider, without affecting inference performance. Enforce token or request-based limits per API key to prevent abuse and protect other users from noisy neighbor effects.”

Buys (vendor) · 3
  • Orb Billing Job post Feb 2026

    “Build and evolve our billing platform and integrations (including Orb), ensuring correctness, auditability, and a high-trust experience for customers and internal teams.”

  • Salesforce CRM Job post 1 Job post 2 Apr 2026

    “Familiarity with Salesforce and how CRM data intersects with comp and territory workflows.”

  • dbt Data platform inferred Job post May 2026

    “Experience building clean, reusable data models and semantic layers in dbt.”

Open roles in the revenue & lifecycle org — 9
View open roles
  • Software Engineer - Billing and Internal Tooling MonetizationBilling engineering seen May 12, 2026
  • Cost Analytics Lead Billing engineering seen May 12, 2026
  • Strategic Finance, GTM Deal desk seen Apr 28, 2026
  • Head of Legal Operations Deal desk seen Apr 28, 2026
  • GTM Engineer RevOps seen Apr 22, 2026
  • Field Operations & Incentives Manager RevOps seen Apr 22, 2026
  • Account Executive - AI Native: Strategic Customer success seen Apr 21, 2026
  • AI Solutions Engineer Customer success seen Apr 21, 2026
  • Applied AI Inference - Forward Deployed Engineer Customer success seen Apr 21, 2026
  • +7 more matched roles

Signals reviewed · derived from public job posts, engineering blogs

Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.

Key takeaways

  1. Per-minute billing is the new transparency floor for inference infrastructure. Baseten’s per-minute GPU rate card made cost-modeling tractable for AI engineering teams in a way per-hour hyperscaler billing never did. Inference platforms targeting cost-conscious AI-native customers should publish per-second or per-minute rates as a baseline expectation.

  2. Open-source tooling is the most efficient PLG channel for infrastructure pricing. Baseten’s open-sourcing of Truss in 2022 made it the de-facto packaging framework and seeded production migration to per-minute hosted deployments. For usage-based infrastructure products, seeding adoption through OSS via guides like our intro to UBP is dramatically more cost-efficient than paid acquisition.

  3. Dual-track usage pricing captures both ends of the workload spectrum. By offering per-minute dedicated AND per-token Model APIs in one workspace, Baseten captures bursty low-volume customers (where per-token wins) and sustained high-QPS customers (where per-minute wins) without forcing a platform choice. Most competitors leave one segment uncaptured by force-fitting workloads to a single SKU model.

  4. BYOC as enterprise upsell preserves the customer relationship past procurement. The self-host option converts the data-residency or committed-spend objection into an Enterprise contract — keeping the customer on Baseten’s control plane and Truss tooling rather than losing them to in-house Kubernetes. This is one of the cleanest enterprise pricing architectures in AI infrastructure.

  5. A hosted-model catalog is a pricing commitment, so publish its exit path too. Baseten’s cached-input column (80–90% below standard input) removed a real switching-cost objection for teams using DeepSeek or Kimi via first-party APIs — but the 2026-07-21 cut showed that a per-model discount evaporates when the model does, and the cheapest cache rate doubled without a single published price moving. Eight days later, on 2026-07-29, the same lever moved favorably instead — GLM-5.2’s cache rate was cut 46% and two pricier flagship SKUs joined the catalog — proving the rate card is a continuously managed surface, not a one-time pruning event. The pattern held again on 2026-08-04, when Baseten grew the catalog to twelve models without retiring anything and, for the first time, renamed an existing SKU (DeepSeek V4 to DeepSeek V4 Pro) rather than repricing or retiring it — a reminder that catalog stability for a usage-based buyer includes identifier stability, not just price stability. A fourth event on 2026-08-28 — two more additions, no retirement or repricing this time — didn’t change that lesson, it confirmed it: the fourth catalog event in five weeks, and the first that touched nothing but additions, still shipped with zero changelog markup on the pricing page itself. Any platform whose rate card is a list of third-party models should treat catalog-change notice as part of the price, because buyers budget against the catalog as it stands today, not just the rates they last checked.


UBP implications

  1. Per-minute granularity is the credible commitment for scale-to-zero economics. Per-hour billing makes scale-to-zero a marketing claim more than a financial reality — rounded-up minimums consume the savings. For usage-based pricing models where idle-state economics are part of the value proposition, billing granularity must match the marketing claim.

  2. Multi-SKU usage billing reduces vendor sprawl but requires unified forecasting. Baseten’s per-minute + per-token dual-SKU model maximizes wallet share within each customer, but customers struggle to forecast mixed-SKU spend. Future usage-based platforms shipping multiple billing dimensions must invest in unified forecasting tools as a first-class product surface, not an afterthought reporting feature.

  3. Per-SKU discounts are only as durable as the SKU — and the same lever cuts both ways. When OpenAI, Anthropic, and now hosted-DeepSeek-via-Baseten all publish 80–90% cached-input discounts, the rate card without caching becomes uncompetitive — but Baseten’s 2026-07-21 catalog cut showed that a discount attached to a line item can be withdrawn by deleting the line item, no price change required. Its 2026-07-29 follow-up showed the inverse: GLM-5.2’s cache rate was cut 46% and two pricier flagship SKUs (Kimi K3, GLM-5.2 Fast) joined the catalog within the same eight-day window, all at Baseten’s sole discretion. A third event on 2026-08-04 surfaced a different risk: Baseten renamed an existing SKU (DeepSeek V4 to DeepSeek V4 Pro) without changing its price or retiring it — a stability risk distinct from pricing, since any customer or downstream billing integration keyed to the model name rather than a stable model ID has no signal that the identifier moved. A fourth event on 2026-08-28 — pure addition, no repricing or rename — left the underlying lesson unchanged: the catalog is a live surface on a weeks-not-quarters cadence, and any vendor pricing a multi-tenant model catalog should expect buyers to re-check it that often. Usage-based vendors should expect buyers to start pricing catalog and identifier stability alongside the rates themselves, and account-level entitlements (a cache discount that applies to whatever you run, addressed by a stable ID rather than a display name) will read as a stronger commitment than a per-model column that can move up, down, or simply be relabeled without notice.


Sources


Bottom line

Baseten priced inference infrastructure for the post-hyperscaler era: per-minute GPU rate cards published openly, no markup on hosted open-source models, scale-to-zero with configurable warm pools, and a dual per-minute + per-token model that lets a single workspace cover both low-volume bursty queries and sustained production workloads. The Truss-based developer experience and BYOC enterprise upsell make the platform sticky once adopted.

For AI-native engineering teams cost-modeling production inference, Baseten is the most legible commercial alternative to building on raw Kubernetes — and the transparent rate card is itself a strategic asset. The remaining gaps (hidden Pro discount schedules, cold-start economics not surfaced on pricing pages, no unified mixed-SKU forecasting, and an undocumented, continuously-churning Model APIs catalog — eleven SKUs to eight on 2026-07-21, eight to ten on 2026-07-29, ten to twelve plus a silent SKU rename on 2026-08-04, twelve to fourteen on 2026-08-28) are GTM polish problems rather than structural pricing flaws — and the dedicated per-minute path, where the customer owns the model, is unaffected by all of them.

Compare with peers via the blueprint corpus, or model your own spend using the Baseten pricing calculator.

Pricing timeline : Major events on a vertical axis

Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.

Model APIs Catalog Grows to Fourteen Models

Three weeks after the DeepSeek V4 Pro rename, Baseten added GLM-5.3-Flash ($0.15 input / $0.03 cache input / $0.50 output per 1M tokens) and DeepSeek V4 Pro 0813 ($1.32 / $0.132 / $3.96 — a dated DeepSeek V4 Pro variant with a 24% lower input rate but 14% higher output rate). No rate on any of the twelve prior SKUs changed, and the homepage banner switched from promoting DeepSeek V4 Flash to promoting GLM-5.3-Flash. Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

Model APIs Catalog Grows to Fourteen Models - Three weeks after the DeepSeek V4 Pro rename, Baseten added GLM-5.3-Flash ($0.15
captured

Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro

Six days after growing to ten models, Baseten added DeepSeek-V4-Flash-0731 ($0.13 input / $0.028 cache input / $0.26 output per 1M tokens — the platform's cheapest model to date) and Inkling-Small ($0.50 / $0.10 / $1.20), and renamed the existing DeepSeek V4 SKU to DeepSeek V4 Pro at unchanged rates ($1.74 / $0.145 / $3.48). The homepage banner switched from an 'Announcing our Series F' funding promo to a DeepSeek V4 Flash product promo. Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

Model APIs Catalog Grows to Twelve Models; DeepSeek V4 Renamed Pro - Six days after growing to ten models, Baseten added DeepSeek-V4-Flash-0731 ($0.1
captured

Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46%

Eight days after pruning its Model APIs catalog to eight models, Baseten added flagship Kimi K3 ($3.00 input / $0.30 cache input / $15.00 output per 1M tokens) and GLM-5.2 Fast ($2.10 / $0.21 / $6.60), and cut GLM-5.2's own cache-input rate 46% to $0.14 per 1M tokens (from $0.26; its $1.40 input and $4.40 output rates held steady). Dedicated GPU/CPU/Training rate cards and the Basic/Pro/Enterprise tier structure were unchanged.

Model APIs Catalog Grows to Ten Models; GLM-5.2 Cache Rate Cut 46% - Eight days after pruning its Model APIs catalog to eight models, Baseten added f
captured

Model APIs Catalog Pruned to Eight Models; Inkling Added

Baseten cut its Model APIs rate card from eleven SKUs to eight, retiring GLM 5.1, GLM 5, Kimi K2.5 and NVIDIA Nemotron 3 Super, and added Inkling at $1.00 input / $0.17 cache input / $4.05 output per 1M tokens. No surviving model's rate changed, but the retirements removed the cheap end of the catalog: the lowest published cache-input rate doubled from $0.06 to $0.12. Dedicated GPU, CPU, Training and Basic/Pro/Enterprise surfaces were unchanged.

Model APIs Catalog Pruned to Eight Models; Inkling Added - Baseten cut its Model APIs rate card from eleven SKUs to eight, retiring GLM 5.1
captured

Cached Input Pricing on Model APIs

Baseten added a dedicated Cache Input column to the Model APIs rate card — prompts re-using prefixes from prior requests within a session window are billed far below the standard input rate. This narrows the gap against first-party caching offerings from OpenAI and Anthropic.

Self-Hosted (BYOC) Deployment Option

Baseten launched a self-hosted deployment option for enterprises requiring data-residency control or compute-cost optimization via their own AWS, GCP, or Azure GPU reservations. Pricing shifts from per-minute usage to a platform license fee plus customer-supplied compute. Quoted as Enterprise-only.

B200 GPU Availability + Mission Critical SLA Tier

Baseten added NVIDIA B200 (180GB) at $0.16633/min as its first Blackwell-generation instance, and formalized a Mission Critical SLA tier for enterprise inference workloads with 99.95% uptime guarantees and 24/7 incident response.

H100 GPU Pricing Published at $0.10833/min

Baseten published transparent per-minute pricing for H100 80GB at $0.10833/min (~$6.50/hour), undercutting hosted inference markups by several major hyperscalers while remaining higher than raw AWS on-demand rates. The transparency was positioned as a competitive differentiator over Bedrock and Vertex AI.

Model APIs (Multi-Tenant) Launched

Baseten introduced Model APIs — multi-tenant per-token endpoints for popular open-weight models (DeepSeek, Llama 3, Mixtral, Whisper). Pricing: per million input/output tokens, comparable to first-party rates. This added a pure pay-per-token SKU to complement the per-minute dedicated deployment offering.

Series C ($75M) — IVP-Led at $825M Valuation

Baseten raised a $75M Series C led by IVP at an ~$825M post-money valuation. Customers disclosed at the round: Writer, Descript, Patreon, Robust Intelligence, Picnic Health. Series C funded the launch of Model APIs and Mission Critical SLA tier.

Model Library Launched

Baseten launched its Model Library — a catalog of pre-deployed open-source models (Stable Diffusion, Whisper, Llama 2, Mistral) that customers could deploy with one click. Pricing: per-minute GPU rate of underlying instance type, no model markup. This established the per-minute, per-instance billing model that remains the core SKU today.

Series B ($20M) — Truss Framework Open-Sourced

Baseten raised a $20M Series B led by Greylock and South Park Commons at a $90M valuation, and open-sourced Truss — its model-packaging framework that became the standard interface for deploying PyTorch and TensorFlow models on Baseten infrastructure. Pricing pivoted to consumption-based GPU minutes.

Baseten Founded

Tuhin Srivastava, Amir Haghighat, and Philip Howes (ex-Gumroad) founded Baseten with a focus on letting data scientists deploy ML models without DevOps. Initial product was a no-code model serving interface.

Trivia
  • · Baseten's $0.10833/minute H100 rate works out to ~$6.50/hour — roughly 1.5–2× AWS on-demand H100 list but Baseten markets the spread as the cost of scale-to-zero plus engineer-free ops.
  • · Truss, Baseten's open-source model-packaging framework, predates the company's Model APIs by three years — Baseten started as a 'bring your weights, we run the serving stack' product before adding hosted multi-tenant model endpoints in 2024.
  • · Baseten's $75M Series C in February 2024 was led by IVP at a $825M post-money valuation; customers cited at the round included Writer, Descript, Patreon, and Robust Intelligence.

Questions & answers

How much does Baseten cost per month?
Baseten has no monthly fee on the Basic tier — you pay only for the GPU minutes your dedicated deployments run plus per-token usage on Model APIs. A small production deployment running an A10G GPU 24/7 would cost roughly $870/month ($0.02012 × 60 × 24 × 30); the same model on a scale-to-zero pattern serving ~10 requests/hour might cost under $100/month.
What GPU types does Baseten support and at what price?
Baseten supports T4 ($0.01052/min), L4 ($0.01414/min), A10G ($0.02012/min), A100 80GB ($0.06667/min), H100 MIG 40GB ($0.0625/min), H100 80GB ($0.10833/min), and B200 180GB ($0.16633/min). All rates are per-minute and billed only while the replica is serving traffic.
Does Baseten have a free tier?
Yes — the Basic tier carries no monthly fee. New accounts also receive free credits for experimentation. Basic users get SOC 2 Type II and HIPAA-compliant deployments, in-app and email support, and access to all instance types — they only pay for actual GPU minutes consumed.
How do Baseten's Model APIs compare to dedicated deployments?
Model APIs are multi-tenant per-token endpoints for a curated fourteen-model catalog (Kimi K3, Kimi K2.6, Kimi K2.7 Code, Inkling-Small, Inkling, GLM-5.2, GLM-5.2 Fast, GLM-5.3-Flash, GLM 4.7, NVIDIA Nemotron 3 Ultra, DeepSeek-V4-Flash-0731, DeepSeek V4 Pro, DeepSeek V4 Pro 0813, GPT OSS 120B) — billed at $0.10–$3.00 per million input tokens and $0.26–$15.00 per million output tokens depending on the model. Dedicated deployments are single-tenant per-minute GPU rentals where you control the model and runtime. Model APIs are cheaper for low-volume use; dedicated deployments are cheaper at sustained high QPS or for proprietary models. The Model APIs catalog is also actively managed — Baseten pruned it from eleven models to eight on 2026-07-21, grew it to ten on 2026-07-29 with flagship Kimi K3 and GLM-5.2 Fast, to twelve on 2026-08-04 with DeepSeek-V4-Flash-0731 and Inkling-Small (renaming DeepSeek V4 to DeepSeek V4 Pro), and to fourteen on 2026-08-28 with GLM-5.3-Flash and a second dated variant, DeepSeek V4 Pro 0813 — so a specific checkpoint you need long-term is safer on a dedicated deployment.
What does Baseten Enterprise / Mission Critical SLA include?
Enterprise tier adds custom SLAs (Mission Critical at 99.95% uptime), self-hosted (BYOC) deployment option in your AWS/GCP/Azure VPC, data-residency controls, advanced RBAC, custom regions, and dedicated Slack/Zoom support. Pricing is quote-based and typically tied to annual usage commitments at volume discounts.
Does Baseten charge for idle replicas?
No — Baseten's pricing is explicitly per active minute, and idle replicas in scale-to-zero state incur no charge. The trade-off is cold-start latency on the first request after idle; the platform offers configurable warm-pool windows for latency-sensitive workloads at the cost of slightly higher idle billing.