Ask
Launch

Hugging Face adds AWS H100, B200 and RTX PRO 6000 to Inference Endpoints

Hugging Face pricing

Hugging Face added AWS NVIDIA H100 ($4.50/hr) plus its first Blackwell GPUs, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr), to Inference Endpoints pricing.

Before

AWS GPU catalog topped out at A100 ($2.50/hr) and H200 ($5/hr); H100 was available only on GCP at $10/hr; no Blackwell-generation hardware existed.

After

AWS now also offers H100 ($4.50/hr), plus two new Blackwell-generation options, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); ZeroGPU Spaces now disclose they run on RTX PRO 6000 Blackwell hardware.

Proof of change

Hugging Face's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jun 15, 2026 Captured Jul 28, 2026 · 43 days apart
main
Captured Jun 15, 2026
Captured Jul 28, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

What these images do and don't show
  • 3 pages had no counterpart in the earlier capture and are not shown.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Hugging Face expanded the GPU catalog behind its per-hour Inference Endpoints product with three new instance types on AWS: NVIDIA H100 ($4.50/hr), which undercuts the existing GCP-hosted H100 rate of $10.00/hr by more than half, and two brand-new Blackwell-generation accelerators — NVIDIA B200 ($9.25/hr) and NVIDIA RTX PRO 6000 ($2.75/hr). All three scale linearly at 2x/4x/8x multiples on the same per-hour basis as the rest of the endpoint catalog.

The change is visible on the main huggingface.co/pricing page as of this capture, but had not yet propagated to the dedicated Inference Endpoints docs pricing table, which still lists only the prior AWS lineup (T4 through H200) and the GCP-only H100. Separately, Hugging Face’s Spaces hardware page now names the underlying accelerator behind the free ZeroGPU tier for the first time: an NVIDIA RTX Pro 6000 Blackwell card with up to 96 GB of VRAM, still free for PRO/Team/Enterprise accounts. Seat subscription pricing (PRO $9/mo, Team $20/user/mo, Enterprise $50/user/mo) and Inference Providers credits are unchanged.

From Hugging Face's pricing timeline
AWS adds H100 and first Blackwell GPUs (B200, RTX PRO 6000) to Inference Endpoints

Hugging Face expanded the AWS instance catalog behind Inference Endpoints with NVIDIA H100 ($4.50/hr, undercutting GCP's $10/hr H100 by more than half) plus its first Blackwell-generation accelerators, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); Spaces' free ZeroGPU tier also disclosed for the first time that it runs on RTX Pro 6000 Blackwell hardware.

About Hugging Face
huggingface.co ↗

Hugging Face monetizes a free model/dataset Hub through four publicly-priced surfaces: seat subscriptions, per-hour inference endpoints, per-hour Spaces hardware, and pass-through per-token inference.

Free tier
Yes
Commits
None
Transparency
public

Hugging Face pricing history

  1. Aug 2026
    Rate card re-verified; Enterprise adds a 'Cost management' row
  2. Aug 2026
    Rate card re-verified; Enterprise gains a granular RBAC row, Baseten joins Inference Providers
  3. Jul 2026
    AWS adds H100 and first Blackwell GPUs (B200, RTX PRO 6000) to Inference Endpoints
  4. Jun 2026
    Current rate card verified
  5. Jul 2025
    Inference Providers replaces the serverless Inference API
Full Hugging Face timeline
All pricing activity