Launch

Hugging Face adds AWS H100, B200 and RTX PRO 6000 to Inference Endpoints

Hugging Face pricing

Hugging Face added AWS NVIDIA H100 ($4.50/hr) plus its first Blackwell GPUs, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr), to Inference Endpoints pricing.

Before

AWS GPU catalog topped out at A100 ($2.50/hr) and H200 ($5/hr); H100 was available only on GCP at $10/hr; no Blackwell-generation hardware existed.

After

AWS now also offers H100 ($4.50/hr), plus two new Blackwell-generation options, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); ZeroGPU Spaces now disclose they run on RTX PRO 6000 Blackwell hardware.

Hugging Face expanded the GPU catalog behind its per-hour Inference Endpoints product with three new instance types on AWS: NVIDIA H100 ($4.50/hr), which undercuts the existing GCP-hosted H100 rate of $10.00/hr by more than half, and two brand-new Blackwell-generation accelerators — NVIDIA B200 ($9.25/hr) and NVIDIA RTX PRO 6000 ($2.75/hr). All three scale linearly at 2x/4x/8x multiples on the same per-hour basis as the rest of the endpoint catalog.

The change is visible on the main huggingface.co/pricing page as of this capture, but had not yet propagated to the dedicated Inference Endpoints docs pricing table, which still lists only the prior AWS lineup (T4 through H200) and the GCP-only H100. Separately, Hugging Face’s Spaces hardware page now names the underlying accelerator behind the free ZeroGPU tier for the first time: an NVIDIA RTX Pro 6000 Blackwell card with up to 96 GB of VRAM, still free for PRO/Team/Enterprise accounts. Seat subscription pricing (PRO $9/mo, Team $20/user/mo, Enterprise $50/user/mo) and Inference Providers credits are unchanged.

From Hugging Face's pricing timeline
AWS adds H100 and first Blackwell GPUs (B200, RTX PRO 6000) to Inference Endpoints

Hugging Face expanded the AWS instance catalog behind Inference Endpoints with NVIDIA H100 ($4.50/hr, undercutting GCP's $10/hr H100 by more than half) plus its first Blackwell-generation accelerators, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); Spaces' free ZeroGPU tier also disclosed for the first time that it runs on RTX Pro 6000 Blackwell hardware.

About Hugging Face
huggingface.co ↗

Hugging Face monetizes a free model/dataset Hub through four publicly-priced surfaces: seat subscriptions, per-hour inference endpoints, per-hour Spaces hardware, and pass-through per-token inference.

Free tier
Yes
Commits
None
Transparency
public

Hugging Face pricing history

  1. Jul 2026
    AWS adds H100 and first Blackwell GPUs (B200, RTX PRO 6000) to Inference Endpoints
  2. Jun 2026
    Current rate card verified
  3. Jul 2025
    Inference Providers replaces the serverless Inference API
  4. Sep 2024
    Enterprise Hub repriced to per-seat with included usage
  5. May 2023
    Inference Endpoints launched; per-hour managed GPUs
Full Hugging Face timeline
All pricing activity