Hugging Face adds AWS H100, B200 and RTX PRO 6000 to Inference Endpoints
Hugging Face added AWS NVIDIA H100 ($4.50/hr) plus its first Blackwell GPUs, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr), to Inference Endpoints pricing.
AWS GPU catalog topped out at A100 ($2.50/hr) and H200 ($5/hr); H100 was available only on GCP at $10/hr; no Blackwell-generation hardware existed.
AWS now also offers H100 ($4.50/hr), plus two new Blackwell-generation options, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); ZeroGPU Spaces now disclose they run on RTX PRO 6000 Blackwell hardware.
Proof of change
Hugging Face's pricing pages, as we captured them on two dates.
Showing the whole page as captured — scroll either panel, or open it at full size.
- 3 pages had no counterpart in the earlier capture and are not shown.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Hugging Face expanded the GPU catalog behind its per-hour Inference Endpoints product with three new instance types on AWS: NVIDIA H100 ($4.50/hr), which undercuts the existing GCP-hosted H100 rate of $10.00/hr by more than half, and two brand-new Blackwell-generation accelerators — NVIDIA B200 ($9.25/hr) and NVIDIA RTX PRO 6000 ($2.75/hr). All three scale linearly at 2x/4x/8x multiples on the same per-hour basis as the rest of the endpoint catalog.
The change is visible on the main huggingface.co/pricing page as of this capture, but had not yet propagated to the dedicated Inference Endpoints docs pricing table, which still lists only the prior AWS lineup (T4 through H200) and the GCP-only H100. Separately, Hugging Face’s Spaces hardware page now names the underlying accelerator behind the free ZeroGPU tier for the first time: an NVIDIA RTX Pro 6000 Blackwell card with up to 96 GB of VRAM, still free for PRO/Team/Enterprise accounts. Seat subscription pricing (PRO $9/mo, Team $20/user/mo, Enterprise $50/user/mo) and Inference Providers credits are unchanged.
Hugging Face expanded the AWS instance catalog behind Inference Endpoints with NVIDIA H100 ($4.50/hr, undercutting GCP's $10/hr H100 by more than half) plus its first Blackwell-generation accelerators, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); Spaces' free ZeroGPU tier also disclosed for the first time that it runs on RTX Pro 6000 Blackwell hardware.