Infrastructure-Layer AI Vendors Standardize on Commitment Pricing
Seventy-nine percent of infrastructure-layer AI companies in the corpus have commitment pricing — reserved capacity, throughput reservations, or volume commitments — versus 33% corpus-wide. GPU capacity economics make commitments a structural necessity at the infra layer.
What's happening — and why
What's happening: fifteen of nineteen infrastructure-layer AI companies (GPU clouds, inference APIs, serving platforms) publish commitment pricing — typically annual reserved capacity, throughput reservations, or volume commits at 20-40% below PAYG. The corpus-wide commitment rate is 33%.
Why: GPU infrastructure requires capacity planning on both sides of the transaction. Providers with fixed silicon costs (Cerebras' wafer-scale chips, Groq's LPUs, dedicated H100 clusters) cannot absorb demand uncertainty. A committed-use contract lets the vendor guarantee utilization; the buyer gets a lower rate and SLA guarantees on throughput or latency — critical for production AI workloads.
Modal's February 2026 addition of AWS/GCP Marketplace billing extends this: enterprise buyers increasingly want to consume AI infra through existing cloud commitments, so providers are adding Marketplace channels where cloud spend counts toward the AI bill.
How it works
Evidence over time
27 supporting · 2 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| anyscale | Jan 2024 | Annual committed-use contracts layered on top of per-Anyscale-Credit PAYG; BYOC enterprise option |
| baseten | Jan 2025 | Dedicated GPU deployment commitments (annual or multi-year) plus per-GPU-minute PAYG for shared |
| browserbase | Jan 2025 | Enterprise browser-hours commitment pricing above the self-serve tiers |
| cerebras | May 2026 | Cerebras Code subscription launched as fixed-price commitment; inference PAYG separate |
| deepinfra | Jan 2025 | Annual committed-use discounts available on inference API; PAYG baseline published |
| e2b | Jan 2025 | Enterprise sandbox-hours commitments on top of compute-credit PAYG |
| fireworks-ai | Jan 2025 | Committed throughput reservations (TPM reservations) available for enterprise; PAYG default |
| groq | Jan 2025 | Volume commit tiers above PAYG; has_commits: true. Throughput reservation for latency SLAs |
| lightning-ai | Jan 2025 | Team/Enterprise plans include committed compute credits; Studio seat + GPU usage hybrid |
| modal | Feb 2026 | AWS and GCP Marketplace billing added for enterprise — cloud committed spend counts toward Modal |
| replicate | Jan 2025 | Enterprise volume commitments available; standard per-second GPU billing as baseline |
| runpod | Jan 2025 | Reserved instance pre-pay discounts vs on-demand; three-tier: on-demand, reserved, spot |
| together-ai | Jan 2025 | Dedicated clusters and throughput reservations for enterprise; public PAYG baseline |
| turbopuffer | Jan 2025 | Monthly minimum floor scales by tier; effectively a soft commitment |
| vast-ai | Jan 2025 | Reserved GPU contracts at discounts vs on-demand spot; three-tier: on-demand, interruptible, reserved |
| lambda-labs | Jun 2026 | Multi-year commitment pricing for B200 GPUs (from $2.99 with multi-year commitment vs $6.69 on-demand) — the widest commit-vs-PAYG spread in the corpus at ~55% off for multi-year reservation. |
| weaviate | Jun 2026 | Vector DB with has_commits: true — enterprise commitment contracts on top of the per-dimension PAYG baseline. |
| pinecone | Jun 2026 | Vector DB with has_commits: true — enterprise annual commitments alongside self-serve per-request/storage billing. |
| milvus | Jun 2026 | Managed vector DB with has_commits: true — commitment pricing for GPU-hours and storage-gb at enterprise tier. |
| bright-data | Jun 2026 | Data infra with has_commits: true — commitment pricing available for bandwidth/IP pools at enterprise scale. |
| runpod | Jul 2026 | The enterprise end formalised: a dedicated runpod.io/enterprise page now names reserved baseline capacity, usage-based burst, committed-use pricing and post-paid billing as four explicit commercial components, alongside contractual SLAs, SOC 2 Type II / HIPAA / GDPR and clusters from 200+ on-demand to 10,000 reserved. Previously Reserved Clusters was a bare 'Contact sales' row with no published commit structure. No enterprise dollar figures published. |
| hyperbolic | Jul 2026 | Commitment moves DOWN-market: Reserved capacity became self-serve in-app at a discounted prepaid $/GPU/hour with terms from ONE WEEK to one month, paid in full up front with no early termination (larger commitments still route through sales); Private Cloud remains the sales-led multi-month-to-multi-year tier, billed separately rather than from credits. In the same window on-demand rates reset upward — H100 SXM $1.50 → $2.89/GPU/hr, H200 $2.40 → $3.49, B200 $3.50 → $5.99 — so the commitment is now the route back to the prior price. |
| hyperbolic | Aug 2026 | The mechanism confirmed a second time at the same vendor: on-demand rose for the THIRD consecutive capture, making the self-serve reserved tier the only route back to a rate that no longer exists. H100 SXM moved $2.89 to $3.19/GPU/hr (+10%) and H200 $3.49 to $3.99 (+14%), with B200 holding at $5.99 — a cumulative +113% on H100 SXM across three captures ($1.50 June, $2.89 July, $3.19 August). The page's own headline moved to match: the 'Affordable compute' banner previously advertised GPUs 'starting at $0.20/GPU/hr' with no bookable card near that price, and now advertises 'starting at $3.19/GPU/hr' — identical to the H100 SXM rate. The 1-week-to-1-month self-serve Reserved ladder introduced on 2026-07-21 is unchanged. |
| fireworks-ai | Aug 2026 | A forward-dated increase across an entire on-demand GPU card — the first repricing of an already-published on-demand SKU since the product launched in January 2024, by Fireworks' own account; every prior on-demand event had been an addition. Effective 2026-09-01: H100/H200 $7.00 to $8.00/hr (+14%), B200 $10.00 to $13.00 (+30%, the steepest on the card), B300 $12.00 to $15.00 (+25%), GB300 $18.00 to $20.00 (+11%). The page frames it as a two-column current-versus-September table rather than an immediate change, which functions as a 20-day notice window for buyers deciding whether to commit at the old rate. It lands three weeks after Fireworks added GB300 288 GB at $18.00/hr and a 1.5x region-restricted deployment premium gated behind Contact Sales. |
| together-ai | Aug 2026 | The reserved SKU got cheaper per token without the price moving. Together's Provisioned Throughput — a fixed tokens-per-minute reservation billed at $0.05 per PTU-minute since its July 2026 launch — raised published per-PTU capacity substantially while holding that sticker flat: MiniMax M3 from 138,840 to 166,667 input TPM/PTU and 694,200 to 833,333 cached (+20% each), and 23,140 to 41,667 output (+80%); GLM-5.2 output from 9,620 to 11,364 (+18%). Kimi K3 joined as a third supported model on the on-page PTU sizing calculator. A commitment product improving on the entitlement axis rather than the price axis, the mirror of how allowances are usually cut. |
| LanceDB | Aug 2026 | Commitment becomes the ONLY paid route, which is the strong form of this trend. LanceDB folded its self-serve "Cloud" tier — serverless, usage-based, explicitly no minimum commitment, self-serve signup — into LanceDB Enterprise, leaving free OSS and a sales-led annual-commit Enterprise sold as Managed or BYOC. No no-commitment paid option is publicly documented. Its corpus taxonomy lost pure-usage, retaining commitment+freemium, and sales_motion dropped self-serve. |
| CoreWeave | Aug 2026 | A commitment gate applied at the SHAPE level rather than the account level. CoreWeave split NVIDIA RTX PRO 6000 Blackwell SE into two configurations: High Memory (1,024 GB RAM) keeps $20.00/hr on-demand and $11.09/hr spot, while a new Standard Memory row (512 GB RAM) is spot-only at $9.56/hr in both North America and Europe, with on-demand and inference both showing "Contact sales". The cheaper configuration is reachable only via spot or a conversation — a commit-or-preempt fork inside one GPU SKU. |
Counterexamples
- bright-data · Jul 2026 — Removed commitment tiers rather than adding them. Web Unlocker and SERP API went from a four-tier committed ladder — PAYG $1.5/1k, $1.3 at $499/mo, $1.1 at $999/mo, $1.0 at $1,999/mo, plus Enterprise — to Free / PAYG / a single $499/mo Scale tier (~380–384k results included, $1.3/1k overage) / Enterprise. The two deepest commit rungs are gone, replaced by a recurring 5,000-results-per-month free tier with no card. Only Scraping Functions kept the old $1.5 → $1.0 ladder. Bright Data was previously cited in this trend's evidence as a commitment vendor.
- coreweave · Jul 2026 — Withdrew a commitment discount at the largest pure-play GPU cloud in the corpus: the footnote reading 'Discounts are available for reserved storage capacity for AI Object Storage and Distributed File Storage' was removed, and the $15,000/mo Direct Connect Virtual 100G tier is now listed 'Not Available' (10G Virtual remains $1,500/mo). Its two new Standard Memory CPU shapes publish spot rates only ($3.94/hr AMD Turin 9655P, $2.86/hr Intel Emerald Rapids 8562Y+) with on-demand marked 'Contact Sales.'
- novita-ai · — — Pure PAYG: per-token inference + per-hour GPU + per-second sandbox with no commit tier published; targets individual developers
- fal-ai · — — Per-output model APIs and per-second GPU compute — no published commitment tier; self-serve only
- deepinfra · — — Publishes volume discount tiers but has_commits is true — technically commits are available; the exception is more nuanced
Trivia
-
Hyperbolic made the commitment the only way back to the old price. On 2026-07-21 it reset on-demand H100 SXM from $1.50 to $2.89/GPU/hr (H200 to $3.49, B200 to $5.99) and in the same window made Reserved capacity self-serve at a discounted prepaid rate — with terms as short as ONE WEEK. A near-doubling of the list rate and a newly self-serve commitment are the same move seen from two sides.
-
Commitment pricing is not spreading monotonically. Bright Data DELETED the two deepest rungs of its committed ladder on 2026-07-14 — $1.1 per 1,000 results at $999/mo and $1.0 at $1,999/mo both gone, leaving a single $499/mo Scale tier — and replaced them with a recurring 5,000-results-per-month free tier requiring no card. CoreWeave removed its reserved-storage discount the same week. Where capacity is abundant, vendors buy acquisition instead of hedging utilisation.
-
The widest commit-versus-PAYG spread in the corpus is Lambda Labs at roughly 55%: B200 GPUs from $2.99/hr on a multi-year commitment against $6.69/hr on-demand. That is more than double the 20-40% band this trend originally described, and it means a buyer signing multi-year at Lambda is effectively paying for 13 months of GPU time out of every 29.
-
Exactly 10 of 353 corpus companies sell through an AWS/Azure/GCP Marketplace — Hugging Face, Milvus, Qdrant, Qodo, SambaNova, Upstash, UsageAI, Vantage, Vast.ai and Weaviate. That is the entire observable channel through which an enterprise's existing hyperscaler commitment can be spent on AI infrastructure, and at under 3% of the corpus it is far smaller than the "consume-through-your-cloud-commit" narrative implies.
For buyers
Model the breakeven between PAYG and committed use before signing. The commit discount (20-40%) is real, but volume floors bite if workloads are unpredictable. For GPU infrastructure, ask: (a) what's the minimum commit, (b) what's the discount vs PAYG, (c) does it count toward existing cloud Marketplace commitments (AWS/GCP/Azure).
For vendors
Commitment pricing at the infra layer is table stakes — buyers expect it once they reach production scale. Design your commitment tier to cover utilization risk: throughput reservations (TPM) for latency-sensitive workloads, reserved capacity (instance reservations) for stable GPU workloads, and Marketplace billing for enterprises with cloud EDPs.
Outlook — what to watch
Cloud Marketplace billing as an enterprise channel will expand — Modal, Anyscale, Groq, Together, RunPod, Replicate, and Baseten already have it. The direction is toward AI infra becoming a line item on existing cloud commits, not a separate vendor contract. Watch for AWS/GCP/Azure adding AI-specific commit categories.
Bottom line
79% of infra-layer AI vendors have commitment tiers — the highest segment rate in the corpus. GPU capacity economics require it on both sides; buyers get 20-40% discounts, vendors get utilization guarantees.
FAQ
Do AI infrastructure vendors offer discounts for commitment?
Yes — 79% of infra-cloud vendors in the corpus (15 of 19) have commitment pricing, with typical discounts of 20-40% over PAYG rates.
What is GPU reserved capacity pricing?
A pre-committed contract for a specific GPU configuration (e.g., 4x A100) for a fixed term (days, months, or a year) at a discounted hourly rate vs on-demand. RunPod, Vast.ai, Together, and Baseten all offer it.
Can I pay for AI infrastructure through my AWS or GCP commitment?
Often yes. Modal, Anyscale, Groq, Together, Replicate, RunPod, and Baseten (among others) offer AWS or GCP Marketplace billing so enterprise spend can draw down existing cloud commits.