Price change

fal splits GPU pricing into list vs 'as low as' rates and adds B300

Fal pricing

fal's Serverless & Compute table now shows two hourly rates per GPU — a List Price and an 'As low as' rate. H100 lists at $3.99/h against the $1.89/h headline. B300 and RTX PRO 6000 join; A100 is gone.

Before

Single per-hour and per-second rate per GPU: A100 40GB $0.99/h, H100 80GB $1.89/h, H200 141GB $2.10/h, B200 184GB 'contact us'.

After

Two published hourly rates per GPU (List Price / As low as): RTX PRO 6000 96GB $2.99/h / $1.10/h, H100 80GB $3.99/h / $1.89/h, H200 141GB $4.50/h / $2.10/h, B200 180GB $6.25/h / $3.49/h, B300 288GB $8.50/h / $4.49/h.

fal rebuilt the compute half of its pricing page. The Serverless & Compute table previously carried one rate per GPU shown at two granularities (per hour and per second); it now carries one granularity (per hour) at two price points — a List Price and an “As low as” rate reserved for custom deployments, with the footer routing buyers to [email protected].

The practical effect is a disclosed discount spread rather than a straight increase. The rates fal used to advertise as the price are now the floor: H100 at $1.89/h and H200 at $2.10/h survive intact in the “As low as” column, while the new List Price column sits roughly 2x higher ($3.99/h and $4.50/h). A self-serve buyer who does not contact sales now sees a materially higher number than they did a month ago.

The fleet itself also moved. B200 is no longer sales-gated — it went from “contact us” to $6.25/h list ($3.49/h as low as), and its VRAM is now listed as 180GB rather than 184GB. B300 (288GB, $8.50/h list, $4.49/h as low as) and RTX PRO 6000 (96GB, $2.99/h list, $1.10/h as low as) are new, and the RTX PRO 6000 replaces the removed A100 40GB as the cheapest entry point. Model API pricing was unchanged in this move: Wan 2.5 $0.05/s, Kling 2.5 Turbo Pro $0.07/s, Veo 3 $0.4/s, Ovi $0.2/video, Seedream V4 $0.03/image, Flux Kontext Pro $0.04/image, Nanobanana $0.0398/image, Qwen $0.02/MP.

From Fal's pricing timeline
GPU compute split into List Price and 'As low as'

The Serverless & Compute table dropped its per-second column and now publishes two hourly rates per GPU. The rates fal used to advertise as the price became the floor (H100 $1.89/h, H200 $2.10/h) while new list prices sit roughly 2x higher (H100 $3.99/h, H200 $4.50/h), so a self-serve buyer who does not contact sales sees a materially higher number. B200 lost its 'contact us' gate at $6.25/h list ($3.49/h as low as); B300 288GB ($8.50/h, $4.49/h) and RTX PRO 6000 96GB ($2.99/h, $1.10/h) joined the fleet; A100 40GB was retired. Model API rates were unchanged.

About Fal
fal.ai ↗

fal (fal.ai) is a generative-media inference platform that prices purely on usage: serverless per-output model APIs plus dedicated GPU compute, with no seats, no subscriptions, and no free tier on its public pricing page.

Pricing model pure usage
Sales motion self servesales led
Free tier
No
Commits
None
Transparency
public

Fal pricing history

  1. Jul 2026
    GPU compute split into List Price and 'As low as'
  2. Jun 2026
    Serverless & Compute + Model APIs
  3. Jul 2025
    Modern layout: Output-Based Pricing + B200 'contact us'
  4. May 2025
    H100 cut to $1.89/hr; full per-hour fleet
  5. Apr 2025
    Per-hour GPU pricing introduced ($1.99/hr H100)
Full Fal timeline
All pricing activity