Hugging Face adds AWS H100, B200 and RTX PRO 6000 to Inference Endpoints
Hugging Face added AWS NVIDIA H100 ($4.50/hr) plus its first Blackwell GPUs, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr), to Inference Endpoints pricing.
AWS GPU catalog topped out at A100 ($2.50/hr) and H200 ($5/hr); H100 was available only on GCP at $10/hr; no Blackwell-generation hardware existed.
AWS now also offers H100 ($4.50/hr), plus two new Blackwell-generation options, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); ZeroGPU Spaces now disclose they run on RTX PRO 6000 Blackwell hardware.
Hugging Face expanded the GPU catalog behind its per-hour Inference Endpoints product with three new instance types on AWS: NVIDIA H100 ($4.50/hr), which undercuts the existing GCP-hosted H100 rate of $10.00/hr by more than half, and two brand-new Blackwell-generation accelerators — NVIDIA B200 ($9.25/hr) and NVIDIA RTX PRO 6000 ($2.75/hr). All three scale linearly at 2x/4x/8x multiples on the same per-hour basis as the rest of the endpoint catalog.
The change is visible on the main huggingface.co/pricing page as of this capture, but had not yet propagated to the dedicated Inference Endpoints docs pricing table, which still lists only the prior AWS lineup (T4 through H200) and the GCP-only H100. Separately, Hugging Face’s Spaces hardware page now names the underlying accelerator behind the free ZeroGPU tier for the first time: an NVIDIA RTX Pro 6000 Blackwell card with up to 96 GB of VRAM, still free for PRO/Team/Enterprise accounts. Seat subscription pricing (PRO $9/mo, Team $20/user/mo, Enterprise $50/user/mo) and Inference Providers credits are unchanged.
Hugging Face expanded the AWS instance catalog behind Inference Endpoints with NVIDIA H100 ($4.50/hr, undercutting GCP's $10/hr H100 by more than half) plus its first Blackwell-generation accelerators, B200 ($9.25/hr) and RTX PRO 6000 ($2.75/hr); Spaces' free ZeroGPU tier also disclosed for the first time that it runs on RTX Pro 6000 Blackwell hardware.