Ask
All companies
technology

Hyperbolic pricing

hyperbolic.ai facts checked analysis reviewed
Quick summary
Region
Product
GPU cloud marketplace & serverless AI inference
Industry
technology
Commits
Available (annual)
In this page
AI Summary
  • Hyperbolic is an open-access AI cloud with four pure usage-based surfaces and no subscription tiers: On-Demand GPUs, Reserved GPUs, Private Cloud, and Serverless Inference.
  • On-demand GPU rental starts at $3.19/GPU/hr (H100 SXM); published starting rates have now risen twice in 2026 — H100 SXM $1.50 (Jun) to $2.89 (Jul) to $3.19 (Aug), H200 now $3.99, B200 $5.99 per GPU-hour — refreshed weekly from supplier rates, with the site's own 'starting at' claim now equal to the H100 rate.
  • Serverless Inference is billed per million tokens, but as of August 2026 the live rate table has narrowed to just two priced SKUs — Llama 3.3 70B and Qwen3-Coder 480B, both $0.40/M tokens — against a broader 'from $0.10/M' text-generation floor claim; Llama-3.1-405B and several other named text, image and vision-language models are flagged for sunset with no priced replacement yet.
  • Everything draws down prepaid compute credits — minimum $5 purchase, 1:1 with dollars, never expiring — with Auto Top-Up recharging the balance below a customer-set threshold.
  • The per-GPU-hour rate shown at instance creation is locked for the life of that instance; reserved instances are prepaid in full at a discount, and Private Cloud is a negotiated custom contract.
  • Account tiers gate request rate and platform access, not token price: the default Free tier (marketed as Basic) is 60 RPM with no minimum spend but cannot launch GPU instances or storage volumes, a one-time $5 deposit auto-upgrades to Pro at 600 RPM, and Enterprise is unlimited and custom-quoted; on-demand and reserved carry a 99.5% uptime SLA and serverless inference 99.9%.
Pricing summary
Hyperbolic 2026 — four pay-as-you-go compute surfaces
Pure usage-based: per-GPU-hour compute in three commitment shapes, plus per-million-token serverless inference. No seats, no subscription.
From $0.10/1M
Serverless Inference
$0.10 /1M tokens
Developers calling open models (Llama, Qwen and more) via an OpenAI-compatible API
Reserved
Discounted $/GPU/hr
Sustained 24/7 workloads that want a lower rate for a fixed term
Private Cloud
Custom contract
Production-scale, single-tenant capacity on multi-month to multi-year commitments
Starting rates as of August 2026 (hyperbolic.ai/marketplace and docs.hyperbolic.ai). GPU rates are refreshed weekly from the best available supplier rates and vary by region and availability — check app.hyperbolic.ai/gpus for live rates before committing. On-demand starting rates ($3.19 H100 SXM, $3.99 H200, $5.99 B200) are unchanged since the prior capture. hyperbolic.ai/inference, the dedicated Serverless Inference marketing page, now returns an HTTP 404 — per-model pricing lives only in the docs, whose live 'Available Models' table shows just two priced SKUs (Llama 3.3 70B and Qwen3-Coder 480B, both $0.40/1M tokens) against broader 'from' claims for text, image and vision-language pricing. Serverless inference is usable on the free account tier (60 RPM); GPU instances and storage volumes require a one-time $5 deposit that promotes the account to Pro.

About

Hyperbolic (Hyperbolic Labs) is a San Francisco-based “open-access AI cloud” founded in 2023 by Jasper Zhang (CEO) and Yuchen Jin (CTO). It now sells compute across four surfaces on a pure pay-as-you-go basis: On-Demand GPUs (hourly, full SSH/root, no commitment), Reserved GPUs (the same hardware prepaid for a fixed term at a discount), Private Cloud (single-tenant, custom contract), and a Serverless Inference API that runs open-weight models (Llama, Qwen, DeepSeek and more) billed per million tokens. The self-serve surfaces are publicly priced with no subscription or seat minimum — buyers purchase compute credits and draw them down by usage.

Hyperbolic’s distinguishing idea is a DePIN-style supply model: rather than build out its own data centers, it aggregates underused GPU capacity from third-party data centers and operators and resells it, which lets it refresh per-hour rates weekly as supply shifts. The company raised roughly $20M total — a seed round around $7M in July 2024 and a $12M Series A in December 2024 led by Variant and Polychain Capital. It reports 200,000+ engineers on the platform and 25+ open-source models served via API, and it is one of Hugging Face’s third-party inference providers. Customers and references include Hugging Face, Quora, Cornell, UC Berkeley, the LMSYS Chatbot Arena, and Reve AI.

For current pricing, see the on-demand and reserved GPU page, the serverless inference docs, and the billing docs. Hyperbolic sits in the AI infrastructure and compute category alongside GPU-cloud rivals and open-model inference hosts.


Pricing summary : GPU-hours, per-million-token inference and prepaid credits

Hyperbolic is pure usage-based across four surfaces with two headline value metrics and a prepaid-credit wallet underneath. There are no subscription tiers, seats, or platform fees — you buy compute credits and every meter draws them down.

  1. On-Demand GPUs — billed per GPU-hour, hourly for the life of the instance. Published starting rates (August 2026): H100 SXM $3.19, H200 $3.99, B200 $5.99 per GPU-hour — the marketplace’s own “starting at” claim now equals the H100 rate, with no separate low-cost floor listed. Rates are refreshed weekly based on the best available rates from suppliers, but the rate displayed at instance creation is locked for that instance’s lifetime. No minimum commitment, no hidden or egress fees.
  2. Reserved — the same GPUs and setup at a discounted prepaid $/GPU/hour, paid in full up front for a fixed term. Self-serve from a 1-week minimum, with longer terms available via sales; no early termination.
  3. Private Cloud — dedicated single-tenant infrastructure on a custom contract, negotiated per deal on multi-month to multi-year commitments and billed separately from credits.
  4. Serverless Inference — billed per million tokens against open-weight models, advertised from $0.10/M for text generation. As of August 2026 the dedicated marketing page (hyperbolic.ai/inference) returns an HTTP 404, and Hyperbolic’s own docs now show a much shorter live rate table: only two Instruct models carry a published per-token price — Llama 3.3 70B and Qwen3-Coder 480B, both $0.40/M tokens with tool calling. Image generation is from $0.0025/image (now shown as an explicit per-pixel/per-step formula) and vision-language models from $0.15/M tokens — but the entire FLUX/Stable-Diffusion image catalog and several named text and vision-language models are flagged “Sunset” in the docs with no replacement models listed yet. Storage volumes attached to instances are metered separately, hourly per GB of provisioned capacity.

Compute credits carry a $5 minimum purchase, are always 1:1 with the dollar amount, and never expire. That $5 is also the account gate: the default Free tier carries no minimum spend but may not launch GPU instances or storage volumes, and a one-time $5 deposit automatically promotes the account to Pro, which unlocks them. To launch an on-demand instance your balance must then cover at least one hour of runtime across all instances; there is no minimum charge beyond that.

What makes this different: two value metrics serve two buyers from one wallet — infrastructure teams that want raw GPU-hours for training and custom serving, and developers that want a per-token serverless API without managing GPUs. And unlike a pure spot market, Hyperbolic pairs weekly, supplier-driven rate refreshes with a per-instance price lock, so the number you saw at launch is the number you keep paying — a hybrid of marketplace pricing and committed-use discounting.


Pricing by product

Hyperbolic sells GPU compute in three commitment shapes plus a serverless token API. Starting rates below are the published on-demand figures as of August 2026; rates are refreshed weekly and vary by region and availability.

GPU compute (commitment shapes)

SurfacePriceIncludedKey mechanics
On-Demand$/GPU/hour, from $3.19/GPU/hrFull SSH / root (VM or bare metal); 8–128+ GPU clusters; up to 24 TB local NVMe; 99.5% uptime SLANo minimum commitment; billed hourly for the life of the instance; rate locked at creation
ReservedDiscounted prepaid $/GPU/hourSame GPUs, networking and setup as On-Demand at a lower rate; 99.5% uptime SLAPaid in full up front; self-serve from a 1-week minimum, longer terms via sales; no early termination
Private CloudCustom contractSingle-tenant, network isolation, optional managed Kubernetes / Slurm; custom SLA; dedicated direct-to-engineer support 24×7Sales-led; multi-month to multi-year commitments, “typically $1M+ deployments”; billed separately, not from credits
Storage volumesHourly, per GB of provisioned capacityNetwork storage that lives independently of instances; a per-GB dollar figure was published in an earlier capture but has since been removed from the platform-comparison docs — no rate is currently listed; docs point to “check current pricing in the console”Billed on total volume size regardless of capacity used or data transferred; terminate anytime

On-Demand GPU starting rates (per GPU-hour)

GPUStarting priceIncludedKey mechanics
NVIDIA B200$5.99 /GPU-hrBlackwell, 192 GB HBM3e, 8 TB/s memory bandwidthFrontier-scale training and highest-throughput inference
NVIDIA H200$3.99 /GPU-hrHopper, 141 GB HBM3e, 4.8 TB/s memory bandwidthLarge-model training and memory-bound inference
NVIDIA H100 SXM$3.19 /GPU-hrHopper, 80 GB HBM3, 3.35 TB/s memory bandwidthMainstream training, fine-tuning and inference
Wider GPU catalog (RTX 4090 / RTX 3080 / RTX 3070)No absolute rate publishedThe provider-comparison table on the marketplace page shows only relative multipliers (e.g. “9.01x cheaper” than AWS for H100 SXM), not $/hr; the on-demand docs list only “H100 80GB, H200 141GB, and B200 192GB”The page’s “starting at” claim now equals the H100 SXM rate ($3.19) — the consumer-GPU catalog floor claim shown in prior months has since been removed entirely; check app.hyperbolic.ai/gpus for live rates

Rates are “refreshed weekly based on the best available rates from suppliers on our platform”, so the per-hour price is dynamic. Node configuration for H100 / H200 is 8 GPUs per node with up to 160 vCPUs, 1.5 TB RAM and 24 TB local NVMe; multi-node clusters use Ethernet or InfiniBand (up to 3.2 Tb/s, NDR / ConnectX-7). Hourly billing carries no hidden or egress fees, and failed instances are never charged.

Serverless Inference (per million tokens)

Model / categoryPriceIncludedKey mechanics
Llama 3.3 70B (Instruct)$0.40 /1M tokens131K context; function/tool calling supportedOne of only two models with a published per-SKU rate as of August 2026
Qwen3-Coder 480B (Instruct)$0.40 /1M tokens262K context; function/tool calling; code-specializedSame rate as Llama 3.3 70B
Text generation (category floor)From $0.10 /1M input tokens”20+” to “25+ open-source models” claimed across the docsFloor claim only — the live “Available Models” table above no longer shows a named small-model SKU at this rate
Image generationFrom $0.0025/image (512×512 @ 25 steps); formula: $0.01 × (width/1024) × (height/1024) × (steps/25)FLUX.1-dev, SDXL 1.0/Turbo, Stable Diffusion 1.5/2, Segmind SD 1BEntire listed catalog is marked “Sunset” — every image model on the docs page is being discontinued with no replacement model named yet
Vision-language modelsFrom $0.15 /1M tokensLlama 3.2 Vision, Qwen2-VLSeveral named VLM SKUs (Qwen2.5-VL-72B/7B-Instruct, Pixtral 12B, NVIDIA Nemotron Nano 12B VL) are separately marked for sunset

As of August 14, 2026 the standalone Serverless Inference marketing page (hyperbolic.ai/inference) still returns an HTTP 404 (first seen 2026-08-11, reconfirmed on each capture since) and is no longer linked from site navigation — per-model pricing now lives only in the docs (inference overview, Text APIs, Image APIs). The live “Available Models” table for chat completions still lists just the two SKUs above, unchanged from the prior capture. The docs’ own “Upcoming Model Deprecations” table newly names five text models being sunset with no priced replacement announced — Qwen3-Next 80B Thinking, Qwen3-Next 80B Instruct, GPT-OSS 120B, GPT-OSS 20B, and Llama 3.1 405B BASE (also flagged individually on the Text APIs page) — alongside the full image catalog and four named vision-language models, so the documented catalog is shrinking on more fronts than the thin 2-SKU live table alone suggests. Audio remains a gap: Melo TTS is sunsetting and Whisper transcription is still “coming soon.” Serverless inference has no minimum commitment and a 99.9% SLA.

Account tiers (rate limits and platform access, not token prices)

TierPricing modelIncludedKey mechanics
Free (marketed as “Basic”)No minimum spend60 RPM rate limit; 100 IP address limit; full precision (BF16) SOTA open-source models; full control over dataDefault plan for new accounts — but “you may not launch GPU instances nor storage volumes with a free-tier account”
ProPay-as-you-go600 RPM rate limit; 100 IP address limit; same model access as Free, plus GPU instances and storage volumes”You must deposit at least $5 one-time to automatically become a pro-tier user” — 10× the request rate at the same per-token price
EnterpriseCustom hourly pricing billed by GPU typeUnlimited rate limit and unlimited IP limit; SOTA open-source models plus custom models; dedicated instances; fine-tuning servicesDedicated support “Available Upon Request”; sales-led — “Contact Us”

Tiers gate request rate and platform access, not the published token price — every tier pays the same per-million-token rate card. The one-time $5 credit deposit is what converts a Free account into a Pro account, so it doubles as the GPU-rental entry ticket. Dedicated model hosting is single-tenant GPU instances with private endpoints, billed hourly with unlimited requests.

Sales motions across products: self-serve / PLG for On-Demand GPUs, Reserved (up to 1 month) and Serverless Inference on the Free/Pro account tiers; sales-led for Private Cloud, large reserved commitments and the Enterprise tier.


Hidden costs : What Hyperbolic users actually pay

Hyperbolic’s headline rates are clean pay-as-you-go, but a few items shape the real bill:

Line itemCost
GPU-hour (e.g. 8× H100 SXM node)$2.89/GPU/hr → ~$23.12/hr for the node
Inference tokensPer-model, $0.10–$4.00 /1M tokens
Storage volumesHourly, per GB of provisioned capacity — billed on total volume size regardless of what you actually use
Credit purchase floor$5 minimum per compute-credit purchase
Launch balance requirementMust hold enough credit to cover 1 hour of runtime across all instances
Weekly rate driftPublished per-hour rate can move week to week for new instances
Reserved / Private CloudPrepaid in full, or sales-quoted per contract

Three real-world cost drivers stand out. First, the headline price is a moving target for new instances: rates refresh weekly off supplier availability, so the rate you budgeted can shift before your next launch — though the price lock means a running instance keeps the rate it started at. Second, storage is billed on provisioned size, not usage — a large volume attached “just in case” bills at full size hourly whether or not you fill it. Third, because supply is aggregated from third parties, capacity and reliability vary by GPU type and region; teams that need guaranteed capacity are steered to Reserved (prepaid, no early termination) or Private Cloud (custom contract). Hyperbolic offsets some of the risk with ready-state billing — instances that never become ready are never charged, and failed reserved instances are refunded in full.

Want to estimate your own Hyperbolic bill? Use the Hyperbolic pricing calculator to model your costs based on GPU type, hours, and token volume.


Pricing evolution : Hyperbolic pricing history and changes

Cadence

PeriodPrice changesProduct / SKU additionsNotes
2024 H2GPU marketplace + serverless inference liveH100 advertised from ~$0.99/hr post-Series A
2025Per-model inference rate card stabilizedImage (SDXL, FLUX), VLM, audio modalities added70B-class at $0.40/M; 405B at $4.00/M
2026 Q2Marketplace starting rates publishedReserved clusters, dedicated hostingH100 from $1.50/hr; weekly supplier-rate refresh
2026 Q35 (Jul: H100, H200, B200 reset upward; Aug: H100, H200 reset again, B200 held)Private Cloud, Reserved self-serve, storage volumes, Auto Top-Up; Reservation Protection documented (not yet live)Jul 21: H100 SXM to $2.89, H200 $3.49, B200 $5.99; consumer-GPU per-hour rates pulled from the page; catalog floor advertised at $0.20/GPU/hr. Aug 4: H100 SXM to $3.19 (+10.4%), H200 to $3.99 (+14.3%), B200 unchanged at $5.99; “starting at” floor claim corrected to $3.19 — now equal to the H100 rate

Tracked range: 2024–present. Published GPU rates are dynamic (weekly supplier-rate refresh), so point-in-time figures reflect the capture date.

Notable changes

  • Late 2024 — After a $12M Series A (Dec 2024, led by Variant and Polychain), both surfaces were live; early marketing advertised H100 rental from roughly $0.99/hr and per-million-token open-model inference.
  • 2025 — Inference settled into a per-model rate card (3B–8B Llama at $0.10/M, 70B-class at $0.40/M, DeepSeek-V2.5 at $2.00/M, Llama-3.1-405B at $4.00/M), with image, VLM, and audio modalities added.
  • June 2026 — Marketplace published on-demand starting rates (H100 SXM $1.50, H200 $2.40, B200 $3.50, RTX 4090 $0.30, RTX 3070 $0.16), refreshed weekly from supplier rates.
  • July 2026 — Published starting rates reset sharply upward: H100 SXM $2.89, H200 $3.49, B200 $5.99 per GPU-hour, with the consumer-GPU per-hour rates removed from the page in favour of a “from $0.20/GPU/hr” catalog floor. Packaging is now four surfaces (On-Demand, Reserved, Private Cloud, Serverless Inference) with prepaid compute credits ($5 minimum, never expiring), Auto Top-Up, per-instance price lock, separately metered storage volumes, and 99.5% / 99.9% uptime SLAs. Serverless per-million-token rates are unchanged.
  • August 4, 2026 — On-demand starting rates reset for the third time in 2026, and the second time in three weeks: H100 SXM $2.89 → $3.19 (+10.4%), H200 $3.49 → $3.99 (+14.3%); B200 held flat at $5.99. The marketplace’s own “starting at” claim was corrected from an unverifiable $0.20/GPU/hr floor to $3.19 — now identical to the published H100 SXM rate. The billing docs also newly document a not-yet-live “Reservation Protection” control and describe Private Cloud as “typically $1M+ deployments”; serverless token rates and account-tier mechanics did not change.

The 2026 rate resets in detail

Between the 15 June and 4 August 2026 captures, Hyperbolic reset its published on-demand starting rates twice — and the second reset landed just two weeks after the first, not five. H100 SXM moved $1.50 → $2.89 → $3.19 (June → July → August), a cumulative +112.7% in seven weeks; H200 moved $2.40 → $3.49 → $3.99 (+66.3% cumulative); B200 moved $3.50 → $5.99 in July and then held flat in August (+71.1% cumulative, unchanged in the latest capture). Nothing about the pricing mechanism changed across either reset — the page still says rates are “refreshed weekly based on the best available rates from suppliers” — which is precisely the point: a marketplace rate that floats with aggregated third-party supply can float upward as hard as it floats down, repeatedly, without a repricing announcement, a migration path, or a grandfathering policy, because none of those artefacts exist in a spot model.

Who actually paid each increase is the more useful question. Because the per-instance price lock holds the rate shown at instance creation for that instance’s life, a team already running an H100 job through the June rate kept $1.50/hr through both resets; the July jump applied only to instances launched after July 21, and the August jump only to instances launched after August 4. That splits the buyer population cleanly on every reset: long-running workloads absorb nothing, while burst and experiment traffic — the buyers most attracted by a $1.50 headline — now face a rate that has more than doubled across two launches in under two months. Notably, the one card that didn’t move in August, B200, is also the card whose premium over H100 has compressed the most: B200 priced at 2.33x the H100 rate in June, 2.07x in July, and 1.88x in August — the resets are narrowing the spread between the cheapest and priciest listed GPU rather than scaling every tier by the same multiple, which reads more like tightening supply pulling the lower tiers up toward a ceiling than a uniform repricing.

The delisting story got a partial fix. RTX 4090 at $0.30 and RTX 3070 at $0.16 were the cheapest verifiable numbers on the page; when July’s reset replaced them with a “from $0.20/GPU/hr” catalog floor with no matching card, it left the budget tier priced but no longer checkable without signing in. On August 4, Hyperbolic corrected that specific claim — the “starting at” figure now reads $3.19/GPU/hr, exactly the published H100 SXM rate — so the headline number is verifiable again. That is a narrower fix than restoring the RTX catalog itself, which remains delisted, but it closes the gap between an advertised number and a checkable one. Serverless inference, meanwhile, did not move a cent across any of the three captures: the $0.10–$4.00 per-million-token card is byte-for-byte identical in June, July and August, the sharpest available evidence that Hyperbolic deliberately runs a volatile meter and a stable meter side by side.

The direction of travel is a maturing four-surface cloud whose GPU rates are resetting on an increasingly tight cadence — five weeks between the June and July resets, two weeks between July and August — while per-instance price locks and a self-serve Reserved ladder starting at one week absorb the shock for buyers willing to commit, and a stable per-token inference rate card sits untouched on top.


What’s unique : Hyperbolic’s distinctive pricing mechanics

1. Two value metrics, one account. Hyperbolic prices GPU-hours on the marketplace and per-million-tokens on serverless inference from a single prepaid balance — serving infra teams and API developers without forcing either into the other’s billing model.

2. Weekly-refreshed, supplier-driven marketplace rates that move both ways. Because it aggregates third-party GPU supply (a DePIN model), Hyperbolic refreshes per-hour rates weekly off the best available supplier prices — a spot price rather than a fixed list price, which is unusual for raw-GPU rental. 2026 has shown what that means in practice twice already: H100 SXM went from $1.50 to $2.89/GPU/hr in five weeks (June→July), then $2.89 to $3.19 just two weeks later (July→August), with no announcement either time, because a marketplace rate has no announcement to make.

3. Published rates plus a documented price lock — and the lock is the load-bearing half. Hyperbolic publishes both per-GPU-hour starting rates and per-model token rates openly (no sales call to see numbers), and then locks the per-hour rate shown at instance creation for that instance’s whole lifetime. In a stable market that reads as a nicety; after two resets in seven weeks (Jun–Aug 2026, +112.7% cumulative on H100 SXM) it is the actual product, because it is the only thing that stopped back-to-back supplier repricings from landing on workloads already in flight. The documented exception is equally telling — Hyperbolic reserves the right to ask for an increase on instances running “at significantly below-market rates for a long time”, so the lock is a shock absorber, not a perpetual grandfather clause.

4. A commitment ladder that starts at one week. Reserved capacity is now self-serve in-app from one week to one month at a discounted prepaid rate, with larger terms and single-tenant Private Cloud going through sales. Most GPU clouds jump straight from hourly to annual contracts; a one-week prepaid step lets a buyer buy predictability for a single training run rather than a fiscal year — which is exactly the escape hatch a weekly-floating on-demand rate creates demand for.


Strengths & weaknesses

StrengthsWeaknesses
Transparent per-GPU-hour and per-token rates, published without a sales callPublished rates can reset hard, and repeatedly — H100 SXM +112.7% cumulative across two resets in seven weeks (Jun→Aug 2026)
Per-instance price lock kept in-flight workloads on their original rate through both the July and August 2026 resetsOnly three GPUs still carry a published per-hour rate; the RTX consumer-GPU catalog stays delisted even though the “starting at” floor claim was corrected to match the H100 rate in August 2026
Serverless token card held steady while GPU rates moved — predictability where developers need itAggregated third-party supply can vary in capacity and reliability by GPU and region
Commitment ladder is self-serve from one week, not just annual contractsReserved is prepaid in full with no early termination; Private Cloud is sales-quoted and billed outside credits
Two value metrics from one never-expiring prepaid balance, with Auto Top-UpThe Free tier is API-only — it cannot launch GPU instances or storage volumes until a one-time $5 deposit promotes the account to Pro, on top of a 1-hour runtime balance floor
OpenAI-compatible API, zero data retention on inferenceStorage bills on provisioned size, and image/audio per-generation rates are absent from the pricing page

Billing UX : prepaid compute credits, price lock and Auto Top-Up

  • Compute credits — the account wallet. Purchasable at any time with a $5 minimum, always 1:1 with the dollar amount purchased, held as a positive balance, and they never expire. Deposit from the dashboard (“Deposit” in the upper-right), then complete checkout.
  • Account tier is set by deposit, not by a plan picker — new accounts land on the Free tier (no minimum spend, 60 RPM, 100 IP addresses) which “may not launch GPU instances nor storage volumes”. Depositing at least $5 one-time automatically promotes the account to Pro (600 RPM, full platform access, pay-as-you-go); Enterprise is custom-quoted with unlimited limits. The per-million-token rate card is identical on all three.
  • Auto Top-Up — configure a balance threshold (“When my balance goes below…”) and a top-up amount (“Automatically add…”) against a stored default payment method. Hyperbolic evaluates the balance every 10 minutes, uses a backoff strategy on failed charges, and warns that it “may miss massive spikes of usage that rapidly deplete your balance within that timeframe”; it recommends a threshold of 72h of usage and a top-up of 24h of usage.
  • Price lock — “the cost per hour per GPU displayed during instance creation is locked in for the duration of the instance”, with one stated exception: Hyperbolic “may reach out about a price increase if you have had on-demand instances running at significantly below-market rates for a long time”.
  • Reservation Protection (coming soon) — a not-yet-live control documented in the billing docs: enabling it will prevent a Reserved instance’s automatic termination at the end of its term, instead converting it to an on-demand instance billed at then-current on-demand rates until terminated.
  • Balance requirements — before creating an on-demand instance your balance must cover at least one hour of runtime across all your instances, including the new one. There is no minimum charge — you may terminate at any time, including before the instance becomes ready.
  • Ready-state billing — billing starts the moment an instance is fully provisioned and SSH-accessible. Anything that fails to become ready inside the ~3-hour provisioning timeout is auto-terminated and never charged; a failed reserved instance is refunded in full to your balance immediately. The marketplace page also promises notification “within 3 minutes if an instance fails”.
  • Four separate meters on one balance — serverless inference (per API call), GPU instances (on-demand hourly / reserved up-front), and storage volumes (hourly, per GB of capacity) all draw down credits. Private Cloud is billed separately under contract, not from credits.
  • Zero-balance consequences, spelled out — at zero or negative balance you may be unable to launch instances, running instances may be terminated, and storage volumes are terminated after a 30-day grace period to recover data.
  • Public list pricing — per-GPU-hour starting rates and the per-model token card are published rather than gated behind a sales call; only Reserved rates, Private Cloud and the Enterprise inference tier are quoted.
  • Payment methods — “pay with wire / ACH upfront or monthly, or pay as you go via credit card / stripe.”

Strategic wins : Why Hyperbolic’s pricing decisions worked

1. Framing the rate card as a marketplace, not a list price

By reselling underused third-party GPU capacity at openly published, weekly-refreshed rates, Hyperbolic got the visibility of a public rate card without the commitment of a list price — a wedge into the price-sensitive AI-research and indie-developer segment. July 2026 proved the second half of that trade: raising H100 SXM 93%, H200 45% and B200 71% in a single refresh required no price-increase notice, no grandfathering scheme and no customer email, because “refreshed weekly from supplier rates” had already set the expectation. Hyperbolic did it again two weeks later — H100 SXM +10.4%, H200 +14.3% on August 4 — with no notice the second time either, because the market framing had already absorbed the first. Very few vendors can reprice that steeply, twice in three weeks, without a trust event; framing the meter as a market is what buys that room. See how AI companies structure pricing.

2. Two value metrics that capture two buyers

Pricing GPU-hours for infra teams and per-million-tokens for API developers from one account lets Hyperbolic monetize both the “I want raw compute” and the “I just want a model endpoint” buyer without making either adopt the other’s mental model. Related: outcome-based pricing trends.

3. A stable token rate card on top of a spot GPU market

Layering a fixed per-model inference rate card over a fluctuating spot GPU marketplace gives developers predictability where they want it (token price) while letting raw compute float with supply. The July and August 2026 captures are the cleanest proof that this is a deliberate split rather than an accident of timing: GPU-hour rates rose 45–93% in July and another 10–14% in August, both times against a $0.10–$4.00 per-million-token card that stayed byte-for-byte identical across all three captures. An application team building on the inference API felt nothing on either date; a team renting bare H100s repriced twice in seven weeks. See choosing the right usage metric.


Areas to improve : Gaps in Hyperbolic’s pricing approach

1. Weekly rate drift is now a budgeting hazard, not a footnote

A rate that refreshes weekly is fine for short bursts, but the June-to-August 2026 moves (H100 SXM $1.50 → $2.89 → $3.19) show the drift can be large enough to invalidate a quarter’s capacity plan between two launches of the same job — and the cadence is tightening, not just the magnitude: five weeks between the June and July resets, two weeks between July and August. The price lock protects instances already running and self-serve Reserved now buys a fixed rate from one week out, which covers a single training run — but neither helps a buyer who needs a forward number for a budget they must approve before any instance exists. Publishing a rolling rate history per GPU, or a short forward quote a buyer can hold for a few days, would close the gap between “transparent today” and “plannable next month”. See bill shock and cost unpredictability.

2. The RTX catalog is still delisted, even after the floor claim was fixed

Pulling the RTX 4090 ($0.30) and RTX 3070 ($0.16) per-hour rates in July 2026 left the marketplace advertising a “from $0.20/GPU/hr” floor with no matching card on the page. Hyperbolic corrected that specific gap on August 4, 2026, resetting the “starting at” claim to $3.19/GPU/hr — exactly the published H100 SXM rate — so the headline number is checkable again. But the fix is narrower than it looks: it repriced the claim rather than restoring the catalog, so the RTX tier itself remains delisted and the marketplace no longer advertises (or substantiates) any budget option below H100. Restoring per-card starting rates for the full catalog, even with an explicit “refreshed weekly” caveat, would close the gap between a corrected headline and a complete one.

3. Unpublished image/audio rates

Inference token rates are published, but per-generation image (SDXL, FLUX) and audio rates are not shown on the pricing page — and the modality tabs that should carry them currently fail to load, so a buyer cannot reach the numbers at all from the public surface. Rendering those rates as plain text alongside the token card, rather than behind interactive tabs, would extend the transparency that benefits the token products and make the rates survivable when the app misbehaves.

4. Capacity and reliability transparency

Because supply is aggregated from third parties, capacity and reliability can vary by GPU and region. Real-time availability and SLA clarity (without a sales call) would reduce the gap between a self-serve rate card and a self-serve experience.


Monetization stack & signals : how Hyperbolic builds & buys its revenue engine

Buys 0 Builds 1 2 open roles

The read — where the monetization investment is going

Hyperbolic builds billing in-house — no monetization vendor is named. A platform engineer role owning billing, metering and usage tracking plus a Head of Finance running billing and reconciliation point to a self-built revenue engine atop its prepaid-credit core.

Stack — build vs buy
Builds in-house · 1
  • In-house billing, metering & invoicing In-house build Job post 1 Job post 2 Jun 2026

    “Familiarity with billing systems, metering, usage tracking, and quota enforcement mechanisms”

Open roles in the revenue & lifecycle org — 2
View open roles
  • Staff Software Engineer - Platform/Infrastructure Billing engineering seen Jun 18, 2026
  • Head of Finance RevOps seen Jun 18, 2026

Signals reviewed · derived from public job posts

Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.

Key takeaways

  1. Hyperbolic is pure usage-based across four surfaces — On-Demand, Reserved and Private Cloud GPUs billed per GPU-hour, plus serverless inference billed per million tokens — with no subscription or seat anywhere in the model. For the underlying model, see the introduction to usage-based pricing.
  2. A spot meter buys pricing freedom, and Hyperbolic used it twice — because rates are framed as weekly supplier refreshes rather than a list price, H100 SXM could move $1.50 → $2.89 → $3.19 (+112.7% cumulative) across two resets in seven weeks (Jun–Aug 2026) without a repricing announcement or a grandfathering scheme either time.
  3. A per-instance price lock is what makes a floating rate sellable — the lock kept running workloads on their pre-reset rate through both resets, so each increase landed only on new launches; without it, a cumulative +112.7% move would have been a churn event rather than a line-item change.
  4. The real frictions are rate drift and supply variability, not headline fees; teams that need a knowable number move to prepaid Reserved capacity (self-serve from one week) or a sales-quoted Private Cloud contract.
  5. Delisting prices is a costlier move than raising them — but the cost is temporary if you fix it fast — pulling the consumer-GPU per-hour rates in July 2026 left a “from $0.20/GPU/hr” claim no buyer could verify; Hyperbolic corrected it within two weeks by resetting the claim to $3.19 (exactly the H100 rate), restoring a checkable headline even though the RTX catalog itself is still gone from the page.

UBP implications

  1. A spot meter and a fixed meter can coexist — and 2026 is the receipt, twice over. Hyperbolic raised GPU-hour rates 45–93% in July 2026 and another 10–14% in August, both times leaving its $0.10–$4.00 per-million-token card untouched, pricing each metric the way its supply behaves. Any business with both volatile and stable cost inputs can run the same split, so long as buyers can tell which meter they are standing on.
  2. Two value metrics widen the addressable buyer set. Offering raw GPU-hours and a per-token API from one account lets a vendor monetize both the infrastructure buyer and the application developer without forcing a single billing model.
  3. Transparency is a wedge, but it is measured on the worst number you publish — and on how fast you fix it. Documenting the price lock, balance requirements and zero-balance consequences in the open lowers buyer friction against rivals that hide pricing behind sales calls — yet the July 2026 update that raised three published rates also removed two, replacing checkable consumer-GPU prices with an unverifiable “from $0.20/GPU/hr” floor. Hyperbolic corrected that specific claim within two weeks (August 4, 2026), resetting it to match the real H100 rate — evidence that an open rate card can repair a credibility gap quickly, but also that it opened one in the first place.

Sources


Bottom line

Hyperbolic is a clean example of pure usage-based pricing for AI compute: GPUs billed per GPU-hour (H100 SXM from $3.19/hr as of August 2026, refreshed weekly off aggregated third-party supply) across three commitment shapes — On-Demand, prepaid Reserved, and contracted Private Cloud — alongside a serverless inference API billed per million tokens ($0.10–$4.00/M across open models, unchanged since 2025). Everything self-serve draws down never-expiring prepaid compute credits, and the rate shown at instance creation is locked for that instance’s life. That lock did real work across 2026’s two rate resets, when the published H100 SXM rate climbed $1.50 → $2.89 → $3.19 (+112.7% cumulative) in seven weeks and each increase landed only on new launches — the clearest illustration of the bargain on offer here, which is a genuinely open rate card in exchange for a rate that can move sharply, and repeatedly, between one job and the next. The other trade-offs are storage billed on provisioned rather than used capacity, an RTX consumer-GPU tier that has stayed delisted from the page since July 2026 even though the “starting at” claim was corrected in August to match the real H100 rate, and supply that varies by GPU and region — which steers reliability-sensitive and budget-bound teams toward Reserved or Private Cloud. Browse the pricing blueprint for more fully-researched company profiles, or compare Hyperbolic against other AI infrastructure and compute companies.

Pricing timeline : Major events on a vertical axis

Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.

Third 2026 rate reset; marketplace floor claim now matches the H100 rate

On-demand starting rates move again: H100 SXM $2.89 to $3.19/GPU-hr (+10.4%), H200 $3.49 to $3.99 (+14.3%); B200 holds at $5.99. The marketplace's 'starting at' claim, previously an unverifiable $0.20/GPU/hr floor with no matching card, now equals the H100 SXM rate exactly. Serverless inference rates and account-tier mechanics are unchanged.

Third 2026 rate reset; marketplace floor claim now matches the H100 rate - On-demand starting rates move again: H100 SXM $2.89 to $3.19/GPU-hr (+10.4%), H2
captured

GPU rates reset upward; four-surface platform with prepaid credits

Published on-demand starting rates move to H100 SXM $2.89, H200 $3.49 and B200 $5.99 per GPU-hour (catalog advertised from $0.20/GPU/hr), and the consumer-GPU per-hour rates are pulled from the page. Packaging is now four surfaces — On-Demand, Reserved, Private Cloud and Serverless Inference — funded by prepaid compute credits (minimum $5, never expiring) with Auto Top-Up, a per-instance price lock, and 99.5%/99.9% uptime SLAs. Serverless token rates are unchanged.

GPU rates reset upward; four-surface platform with prepaid credits - Published on-demand starting rates move to H100 SXM $2.89, H200 $3.49 and B200 $
captured

Marketplace starting rates published; H100 from $1.50/hr

GPU marketplace shows on-demand starting rates: H100 SXM $1.50, H200 $2.40, B200 $3.50, RTX 4090 $0.30, RTX 3070 $0.16 per GPU-hour, refreshed weekly from supplier rates. Inference rate card unchanged.

Per-million-token inference rate card stabilizes

Inference settled into a per-model rate card: $0.10/M for 3B-8B Llama, $0.40/M for 70B-class models, $2.00/M for DeepSeek-V2.5, and $4.00/M for Llama-3.1-405B, with image (SDXL, FLUX) and audio modalities added.

Two usage surfaces live; H100 advertised from ~$0.99/hr

After its $12M Series A (Dec 2024), Hyperbolic offered both a GPU marketplace and a serverless inference API. Early marketing cited H100 rental from roughly $0.99/hr, with open-model inference billed per million tokens.

Trivia
  • · Hyperbolic prices two different value metrics from one account: GPU-hours on its marketplace and per-million-tokens on serverless inference.
  • · It runs a DePIN-style model — aggregating underused GPU capacity from third-party data centers and operators — and accepts wire/ACH upfront or monthly as well as pay-as-you-go card payments.
  • · Marketplace GPU rates are refreshed weekly based on the best available supplier rates, but the per-hour rate shown when you create an instance is then locked for that instance's whole lifetime.

Questions & answers

What is Hyperbolic's pricing model?
Pure usage-based, pay-as-you-go. Hyperbolic bills per GPU-hour on On-Demand and Reserved GPUs and per million tokens on its serverless inference API, with storage volumes billed hourly per GB. There are no subscription tiers or seat fees — you buy compute credits and draw them down by consumption. Private Cloud is a negotiated custom contract billed separately.
How much does an H100 cost on Hyperbolic?
As of August 2026, an NVIDIA H100 SXM starts at $3.19/GPU/hr, H200 at $3.99, and B200 at $5.99 per GPU-hour — and the marketplace's own 'starting at' claim now equals the H100 rate, with no separate low-cost floor listed. Rates are refreshed weekly based on the best available supplier rates, so the per-hour price is dynamic: the published H100 SXM rate was $1.50/GPU/hr in June 2026 and $2.89 in July, so it has risen twice in three months. The rate displayed when you create an instance is locked for that instance's lifetime, so a refresh only affects your next launch.
How is Hyperbolic's inference pricing charged?
Serverless inference is billed per million tokens, with each open-weight model carrying its own rate: $0.10/M for Llama-3.1-8B and Llama-3.2-3B, $0.20/M for Qwen2.5-Coder-32B, $0.40/M for 70B-class models like Llama-3.1-70B and Qwen2.5-72B, $2.00/M for DeepSeek-V2.5, and $4.00/M for Llama-3.1-405B. You pay only for the tokens you consume.
Does Hyperbolic have a free tier?
Yes, but it is API-only. Hyperbolic's account docs define a Free tier as the default plan for new accounts with no minimum spend required, capped at 60 requests per minute and 100 IP addresses — the same limits the inference page markets as its Basic tier. Free-tier accounts may not launch GPU instances or storage volumes; a one-time $5 deposit automatically upgrades you to Pro (600 RPM) and unlocks them. Compute credits are 1:1 with the dollar amount and never expire, and there is no published free token allowance.
What is Hyperbolic's price lock?
The cost per hour per GPU displayed during instance creation is locked in for the duration of that instance. Hyperbolic reserves one exception: in rare circumstances it may reach out about a price increase if you have had on-demand instances running at significantly below-market rates for a long time.