AI Summary
About
fal (legal entity “features and labels”, at fal.ai) is a generative-media inference platform founded in 2021 by Burkay Gur and Gorkem Yurtseven and headquartered in San Francisco. It runs a GPU fleet optimised for fast, low-latency inference of generative models — image, video, audio, and text-to-speech — and exposes them both as ready-to-call model APIs and as dedicated compute customers can deploy their own apps onto.
fal sells inference wholesale to other AI products as much as to individual developers. Its own marketing cites powering roughly 50% of Poe’s image and video generation, low-latency TTS for PlayAI’s voice agents, and multimodal inference for Genspark’s Super Agent. The company positions on two axes competitors struggle to combine: raw speed (custom inference and training kernels) and cost efficiency (a GPU fleet advertised “as low as $1.89/hr for H100”).
That positioning has attracted aggressive funding. fal raised a $9M seed (a16z), a $14M Series A (Kindred Ventures), a $49M Series B in February 2025 (Notable Capital + a16z), a $125M Series C in July 2025 at a $1.5B valuation (Meritech), and a $140M Series D in December 2025 at a $4.5B valuation led by Sequoia with NVIDIA’s NVentures participating — roughly tripling its valuation in five months on revenue that grew about 1,040% to ~$285M ARR by the end of 2025. The pricing page rebuilt itself alongside that growth, shifting from raw per-second compute billing to today’s per-output model APIs.
Its closest comparison set is other generative-media inference providers and serverless GPU clouds — Replicate, Runpod, and the model-hosting tiers of the hyperscalers. Unlike seat-priced creative SaaS, fal’s entire public pricing surface is metered: you pay per output for model APIs and per hour for GPU compute. There is no free tier, no subscription, and no per-user fee on the published pricing page.
Pricing summary : How fal’s pure-usage model-inference pricing works
fal uses a pure usage-based pricing model — a self-serve, pay-as-you-go motion built on two metered surfaces, with no seats and no subscription:
- Serverless & Compute (GPU rental): Deploy your own app on fal’s GPU fleet, billed per hour at two published rates — a List Price and an “As low as” rate for custom deployments. RTX PRO 6000 96GB is $2.99/h (as low as $1.10/h), H100 80GB $4.50/h (as low as $1.89/h), H200 141GB $4.50/h (as low as $2.10/h), B200 180GB $6.25/h (as low as $3.49/h), and B300 288GB $8.50/h (as low as $4.49/h). H100 and H200 now carry the same $4.50/h list price, though their “As low as” floors still differ. The page says to contact [email protected] to get started; it does not publish the qualifying conditions for the lower rate.
- Model APIs (per output): Call hosted models and pay by output unit, the way individual developers prefer to buy. Video models bill per second or per video (Wan 2.5 $0.05/s, Kling 2.5 Turbo Pro $0.07/s, Veo 3 $0.4/s, Ovi $0.2/video). Image models bill per image or per megapixel (Seedream V4 $0.03/image, Flux Kontext Pro $0.04/image, Nanobanana $0.0398/image, Qwen $0.02/MP). Those summary-table rates are entry rates, not necessarily the rate every call pays: fal’s own model pages price Wan 2.5 at $0.05/second only at 480p ($0.10/second at 720p, $0.15/second at 1080p). As of 2026-08-13, the “Veo 3” row is served by Veo 3.1 — fal deprecated the original fal-ai/veo3 and fal-ai/veo3/fast endpoints and launched Veo 3.1 in their place, priced at $0.20/second with audio off / $0.40/second with audio on for Standard (720p/1080p; $0.40/$0.60 at 4K), and $0.10/second / $0.15/second for Fast (720p/1080p; $0.30/$0.35 at 4K) — roughly 47–60% below the old Veo 3 rates. The pricing table’s $0.4/s now matches Veo 3.1 Standard with audio on.
- Enterprise: Private model hosting, custom kernels, dedicated serverless infrastructure, SLA, SOC2, and SSO — all “Contact Sales,” no public price.
What makes this different: fal publishes prices normalised to “output per $1” (e.g. 20 seconds of Wan 2.5 video or 33 Seedream V4 images), turning the pricing page itself into a buyer-facing cost comparator rather than a plan picker.
Pricing by product
GPU compute (Serverless & Compute)
Deploy your own app on fal’s GPU fleet. Every GPU carries two published hourly rates — a List Price and an “As low as” rate for custom deployments; contact [email protected] to get started. The hero copy still leads with “as low as $1.89/hr for H100.”
| GPU | VRAM | List Price | As low as | Key mechanics |
|---|---|---|---|---|
| RTX PRO 6000 | 96GB | $2.99/h | $1.10/h | Cheapest listed GPU; workstation-class Blackwell |
| H100 | 80GB | $4.50/h | $1.89/h | Headline “as low as $1.89/hr” GPU; list price now matches H200 |
| H200 | 141GB | $4.50/h | $2.10/h | Larger VRAM for bigger models |
| B200 | 180GB | $6.25/h | $3.49/h | Blackwell capacity, now publicly priced |
| B300 | 288GB | $8.50/h | $4.49/h | Largest VRAM on the fleet; top of the list-price range |
Model APIs — Video models
Video models are billed by output unit — per second or per video — depending on the model.
| Model | Unit | Price | Output per $1 | Key mechanics |
|---|---|---|---|---|
| Wan 2.5 | second | $0.05 | 20 seconds | Cheapest listed; $0.05/s is the 480p rate ($0.10/s 720p, $0.15/s 1080p) |
| Kling 2.5 Turbo Pro | second | $0.07 | 14 seconds | Mid-tier per-second video ($0.35 for a 5s clip) |
| Veo 3 | second | $0.4 | 3 seconds | fal.ai/pricing still labels this row “Veo 3”; the model itself is Veo 3.1 as of 2026-08-13 (see caveat below) |
| Ovi | video | $0.2 | 5 videos | Billed per whole video, not per second |
Page footnotes: “For a fair comparison, we’ve normalized these values to show approximate output per $1” and “Based on an estimated average video = 5 seconds at 720p. Actual output may vary based on model, resolution, and prompt complexity.”
Resolution and variant caveat (from fal’s per-model pages, not the summary table): Wan 2.5 is $0.05/second at 480p, $0.10/second at 720p, and $0.15/second at 1080p — so at the 720p basis the table’s own footnote assumes, a dollar buys 10 seconds, not 20. Kling 2.5 Turbo Pro’s model page confirms the $0.07/s rate (“For 5s video your request will cost $0.35”).
Veo repricing (2026-08-13): fal deprecated the original fal-ai/veo3 and fal-ai/veo3/fast endpoints — each now shows “This endpoint is deprecated — This model is no longer supported” — and launched Veo 3.1 (fal-ai/veo3.1) in their place, adding true 4K output, image-to-video, first/last-frame, reference-to-video, and an extend-video mode (chains clips up to ~148s). Veo 3.1 Standard is $0.20/second (audio off) / $0.40/second (audio on) at 720p/1080p, and $0.40/second / $0.60/second at 4K. Veo 3.1 Fast is $0.10/second (audio off) / $0.15/second (audio on) at 720p/1080p, and $0.30/second / $0.35/second at 4K. These are 47–60% below the prior Veo 3 rates ($0.50/$0.75 Standard, $0.25/$0.40 Fast). The summary table’s $0.4/s now matches Veo 3.1 Standard with audio on at 720p/1080p — a coincidental match, since the same figure previously matched the old Veo 3 Fast-with-audio rate.
Model APIs — Image models
Image models are billed by image count or by output size in megapixels (MP); fal normalises listed prices to 1MP for comparison.
| Model | Unit | Price | Output per $1 | Key mechanics |
|---|---|---|---|---|
| Qwen | megapixel | $0.02 | 50 megapixels | Cheapest, billed per MP |
| Seedream V4 | image | $0.03 | 33 images | Per-image billing |
| Flux Kontext Pro | image | $0.04 | 25 images | Per-image billing |
| Nanobanana | image | $0.0398 | 25 images | Per-image billing |
Page footnote: “Output is based on 1MP images. Higher resolutions will be priced proportionally. All models listed here follow output-based pricing. Some other models may use GPU-based pricing depending on architecture.” — i.e. the listed per-image rates only hold at 1MP, and models outside this table can fall back to GPU-time billing.
fal’s per-model pages confirm Seedream V4 at $0.03 per image, Flux Kontext Pro at $0.04 per image, and Qwen at $0.02 per megapixel (rounded up to the nearest megapixel). The Nano Banana model page quotes $0.039 per image where the summary table lists $0.0398 — a rounding difference on the same model, so treat the summary figure as approximate at the fourth decimal.
Enterprise (separate sales-led tier)
| Tier | Price | Included | Key mechanics |
|---|---|---|---|
| Enterprise | Contact sales | Private model hosting, custom inference/training kernels, foundational model research, dedicated serverless infra, SLA, SOC2, SSO, user management | Sales-led, quoted |
Sales motions across products: self-serve / pay-as-you-go for GPU compute and model APIs at list price; sales-led for the “As low as” custom-deployment rates (via [email protected]) and for enterprise (private hosting, custom kernels, SLA, SOC2, SSO).
Hidden costs : What metered inference actually costs at volume
fal’s per-output rates look tiny in isolation, but generative-media workloads multiply them fast. Two representative archetypes:
A consumer AI app generating short videos
A product shipping 50,000 short clips a month (5-second clips at $0.05/s on a Wan-2.5-class model at 480p), plus 200,000 preview images at $0.03 each on Seedream V4.
| Line item | Monthly cost |
|---|---|
| 50,000 clips × 5s × $0.05/s | $12,500 |
| 200,000 preview images × $0.03 | $6,000 |
| Total | $18,500 |
At this volume the per-second video charge dominates, and two switches nobody thinks of as a pricing decision each multiply the bill. Rendering the same clips at 720p instead of 480p doubles Wan 2.5 to $0.10/s and the video line to $25,000. Swapping the model to Veo 3 at the pricing table’s $0.4/s multiplies the video line roughly 8×, to ~$100,000/mo — and if the call actually lands on standard Veo 3 with audio at $0.75/second rather than Veo 3 Fast, it is 15×, ~$187,500/mo.
A team self-hosting on dedicated GPUs
A team running two H100s continuously for a custom pipeline rather than calling model APIs.
| Line item | Monthly cost |
|---|---|
| 2 × H100 × $1.89/hr × 730 hrs | $2,759 |
| Overflow burst: 1 × H200 × $2.10/hr × 100h | $210 |
| Total | $2,969 |
Self-hosting trades the convenience of per-output API pricing for flat GPU-hour cost — cheaper at high, steady utilisation, but you pay for idle time. The crossover depends entirely on how busy the GPUs stay. Note the rates above are the “As low as” floor fal reserves for custom deployments; a team that signs up self-serve at the List Price pays roughly twice this for the same two H100s, so which column applies to you is worth settling before you model the trade-off.
Want to estimate your own fal bill? Use the fal pricing calculator to model your monthly cost across per-output model-API calls and GPU-hour compute.
Pricing evolution : From GPU rental to normalised per-output model APIs
fal’s pricing model has rebuilt itself three times in under three years: from raw per-second unit billing (2024) to a per-output model comparator (2025) to a two-rate compute card (2026), tracking the company’s pivot from a generic serverless-GPU host into a generative-media inference platform — and four funding rounds that took it from a $49M Series B to a $4.5B valuation.
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2024 Q1 | 0 | 0 | Raw per-second unit pricing: CPU $0.00003/s, A100 $0.001/s, A10G $0.0002/s, T4 $0.00009/s; “Choose your Machine Type” picker. |
| 2024 Q2 | 1 | 0 | 2024-05 — “Choose a budget” comparator UI: GPU rates re-expressed as inferences-per-$20 (SDXL, Whisper); A100 now $0.00111/s, A6000 $0.000575/s. |
| 2025 Q1 | 1 | 2 | 2025-01 H100 80GB ($0.00125/s) added + “Billing Based on Model Output” table (FLUX.1, SD3, Stable Video); 2025-02 $49M Series B banner + hero redesign. |
| 2025 Q2 | 2 | 1 | 2025-04 per-hour GPU display (“H100s from $1.99/hr”); 2025-05 H100 cut to $1.89/h, H200 $2.10/h, A100 $0.99/h, A6000 $0.60/h added. |
| 2025 Q3 | 0 | 1 | 2025-07 modern layout reached (coincident with $125M Series C, $1.5B valuation): “Output-Based Pricing” split into Video/Image Models; B200 “contact us” added. |
| 2026 Q2 | 0 | 0 | 2026-06-01 — section labels now “Serverless & Compute Pricing” and “Model APIs Pricing”; H100/H200/A100 rates unchanged since 2025-05. |
| 2026 Q3 | 1 | 2 | 2026-07-23 — compute table split into List Price / “As low as” columns (H100 $3.99/h vs $1.89/h); per-second column withdrawn; B200 un-gated at $6.25/h; B300 288GB and RTX PRO 6000 96GB added; A100 40GB retired; model API rates unchanged. |
Tracked range: 2024-02–2026-07-23. Quarters not listed above were verified stable (0 price changes, 0 SKU additions).
Notable changes
- 2024-02 — Earliest captured pricing: raw per-second unit rates (CPU/Memory/GPU/Storage) with a machine-type picker; no output-based pricing.
- 2024-05 — Budget-slider comparator introduced; GPU-second rates re-expressed as “inferences per $20” — fal’s first output-normalisation.
- 2025-01 — H100 added; per-output “Billing Based on Model Output” table launched alongside the GPU comparator; new diamond wordmark.
- 2025-02 — $49M Series B (Notable Capital + a16z) announced via on-site banner; pricing hero redesigned.
- 2025-04 — Per-hour GPU pricing introduced, led by “H100s from as low as $1.99/hr.”
- 2025-05 — H100 cut to $1.89/h; full per-hour fleet (H200 $2.10/h, A100 $0.99/h, A6000 $0.60/h) — the rates still live a year later.
- 2025-07 — Modern “Output-Based Pricing” layout (Video + Image Models, “Output per $1”); B200 184GB added as sales-gated “contact us.” Coincided with the $125M Series C at a $1.5B valuation.
- 2026-07-23 — Compute repriced into two published columns, List Price and “As low as.” The rates fal had advertised for 14 months became the floor (H100 $1.89/h, H200 $2.10/h) while list prices landed roughly 2x higher (H100 $3.99/h, H200 $4.50/h). The per-second GPU column was withdrawn, B200 lost its “contact us” gate at $6.25/h, B300 288GB ($8.50/h) and RTX PRO 6000 96GB ($2.99/h) joined, and A100 40GB was retired. Model API rates were unchanged.
The 2024-to-2025 repricing in detail
fal’s pricing history is really the story of a positioning change. In 2024 the page sold generic serverless GPU compute — you picked a machine type and paid per unit-second of CPU, memory, and GPU, the same way you’d reason about a cloud VM. The mid-2024 “budget slider” was the first hint of the eventual strategy: it re-expressed those raw GPU-second rates as “how many SDXL images can $20 buy,” teaching buyers to think in outputs rather than seconds.
The decisive break came across 2025 Q1–Q3, bracketed by funding. The $49M Series B (TechCrunch, Sep 2024 seed coverage; Series B announced Feb 2025) financed the pivot to “the future of AI video,” and the pricing page followed: a dedicated “Billing Based on Model Output” table (Jan 2025) introduced true per-output rates, the GPU fleet moved to a per-hour display (Apr 2025) with H100 cut from $1.99 to $1.89/hr (May 2025), and by the $125M Series C ($1.5B valuation, July 2025) the page had split cleanly into GPU rental and per-output Model APIs with the “Output per $1” comparator. fal went on to raise a $140M Series D in December 2025 led by Sequoia at a $4.5B valuation — roughly tripling its Series C mark in five months — on the back of revenue that grew ~1,040% to ~$285M ARR by the end of 2025.
That looked like a stable end-state for a year — and then the compute half moved again. On 2026-07-23 fal stopped publishing one rate per GPU and started publishing two: a List Price and an “As low as” rate for custom deployments. Nothing was taken away from the buyer who negotiates — H100 still bottoms out at $1.89/h, H200 at $2.10/h — but the buyer who does not now sees roughly double that, and the per-second column that let bursty workloads reason in seconds is gone. The direction of travel is worth naming: fal’s model-API surface, where price discovery happens by comparison shopping, stayed frozen and fully public, while the compute surface, where price discovery happens in a conversation, grew a disclosed discount band. Transparency did not retreat so much as change shape — the sales gate moved off the hardware (B200 and B300 are now openly priced) and onto the rate.
What’s unique : Output-normalised, seatless inference pricing
1. The pricing page is a cost comparator, not a plan picker. fal publishes every model rate alongside a normalised “output per $1” figure — 20 seconds of Wan 2.5 video, 33 Seedream V4 images, 50 megapixels of Qwen. Buyers compare cost-per-output directly across models instead of decoding tiers.
2. Two billing surfaces, one metered philosophy. fal lets you either call hosted models per output (zero ops, pay per second/image) or rent the underlying GPU per hour/second and deploy your own app. The same usage-based logic spans both — there is no seat or subscription anywhere on the public page.
3. Two prices per GPU, both printed on the page. Since 2026-07-23 the compute table shows a List Price next to an “As low as” rate for custom deployments — H100 at $3.99/h or $1.89/h, RTX PRO 6000 at $2.99/h or $1.10/h. Most infrastructure vendors publish one number and negotiate the discount privately; fal publishes the whole spread and withholds only the qualifying conditions, so a buyer knows exactly how much room there is before opening the conversation.
4. The sales gate moved from capacity to rate. Through mid-2026 the newest Blackwell SKU was the thing you had to call about (B200 was “contact us”). Now the entire fleet is publicly priced — B200 $6.25/h, B300 288GB $8.50/h — and what you call about is the discount. fal traded a gate on which hardware you can see for a gate on what you’ll actually pay, which is the more valuable secret to keep.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Fully transparent per-output and per-GPU-hour pricing | No free tier to lower trial friction |
| ”Output per $1” normalisation makes model cost comparison easy | Premium models (Veo 3 at $0.4/s) make bills volatile at scale |
| Whole GPU fleet is publicly priced — Blackwell B200 and B300 included since 2026-07 | The hero’s “as low as $1.89/hr” is now the floor, not the self-serve price: H100 lists at $3.99/h |
| Publishing the “As low as” rate makes the negotiable spread visible before you call | Qualifying conditions for the lower rate are unpublished, and the per-second compute display was withdrawn in 2026-07 |
| Two surfaces (managed APIs vs raw GPU) fit different ops appetites | Choosing between them is left to the buyer — the page gives no break-even guidance |
| No seats — cost scales with usage, not headcount | Bill is fully variable; no predictable monthly floor, no published volume commits |
Billing UX : Named controls on the public pricing surface
- “Output per $1” comparator column — every model-API row shows normalised output (seconds, images, or megapixels) per dollar so buyers can compare cost-efficiency without doing the math.
- “List Price” vs “As low as” GPU columns — the Serverless & Compute table shows each GPU’s headline rate next to the best available custom-deployment rate, making the negotiable spread explicit on the page (e.g. H100 $4.50/h vs $1.89/h).
- “Start Building” vs “Contact Sales” CTAs — every pricing block (Serverless & Compute, Video Models, Image Models) pairs a self-serve entry point with a sales path, separating PLG from sales-led motions inline.
- Enterprise Contact Form with company-size selector — the enterprise capture exposes a structured intake (Just me → 2,000+ people, plus company, phone, and project description) that routes high-volume buyers to sales.
[email protected]as the compute sales channel — the GPU table’s footer routes custom deployments to a support address rather than a quote flow: “Competitive pricing for custom deployments — Contact [email protected] to get started.”- Normalisation footnotes under each table — the video table discloses its “estimated average video = 5 seconds at 720p” assumption and the image table its 1MP basis, so the “Output per $1” numbers carry their own caveats inline.
- Per-model price line on every model page — each model’s own page states the rate in a sentence before you call it (“Your request will cost $0.039 per image”; “For 5s video your request will cost $0.35”), and that page — not the summary table — carries the resolution and variant splits (Wan 2.5 $0.05/$0.10/$0.15 per second at 480p/720p/1080p; Veo 3.1 Standard $0.20/s audio off, $0.40/s audio on at 720p/1080p vs Fast $0.10/$0.15). Check the model page before budgeting from the pricing-page table.
- “This endpoint is deprecated” banner on retired model pages — as of 2026-08-13, the original fal-ai/veo3 and fal-ai/veo3/fast pages carry a deprecation notice (“This model is no longer supported”) plus a “Contact sales for competitive API pricing” CTA, while still quoting a live per-second rate that now matches the replacement Veo 3.1 model rather than their own stale Readme copy — a buyer has to read past the deprecation banner to find the number that actually bills.
Strategic wins : Why fal’s seatless, output-normalised pricing works
1. Pricing the unit the customer actually buys
fal charges per generated output — a second of video, an image, a megapixel — which is exactly the unit a generative-media product produces. This is textbook value-metric alignment: the bill moves with the customer’s own output volume, so cost feels fair and scales with their success rather than their headcount.
2. Turning the price table into a comparator
By normalising every rate to “output per $1,” fal removes the cognitive tax of comparing $0.05/second against $0.03/image. This is a subtle usage-based pricing UX win: the pricing page does the buyer’s cost-modelling for them, lowering the barrier to choosing a model.
3. Two surfaces capture two buyer mindsets
Offering both managed per-output APIs and raw per-hour GPU rental lets fal monetise the convenience-seeker and the cost-optimiser without forcing either into the wrong model — a flexible answer to the AI infrastructure cost question of “rent capacity or pay per result.” It is the same dual-surface bet Replicate and Runpod make, but fal leans harder on the per-output side for generative media.
4. Publishing the discount instead of hiding it
The 2026-07-23 split into List Price and “As low as” solves a problem every infrastructure vendor has: a single public rate either leaves margin on the table with price-insensitive buyers or scares off the ones who need a discount to sign. By printing both ends of the band, fal anchors high, keeps its old headline rate credible as a floor, and turns “contact sales” from a wall into a quantified invitation — the buyer already knows a ~2x reduction is on offer before they email. The cost is real, though: self-serve buyers who never make that call now see double the number they saw in June, which is a predictability problem as much as a pricing one.
Areas to improve : Predictability and trial gaps in a fully-metered model
1. No free tier or trial credit
With every line item metered and no free allowance, a curious developer must commit a payment method before generating a single output. A small monthly free-output grant (a few free images/seconds) would lower trial friction the way credit grants do in other usage-based pricing models, without materially denting revenue.
2. Fully variable bills invite bill shock
Because there is no floor and premium models cost 8–15× the cheapest listed rate (Veo 3 runs $0.4/s on the summary table and $0.75/s with audio on its own model page, against Wan 2.5’s $0.05/s at 480p), a model swap, a resolution bump, or a traffic spike can multiply a bill overnight — the classic AI cost-unpredictability problem. Published spend caps, budget alerts, or committed-use discounts would give finance teams the predictability the current page lacks.
3. The “As low as” rate has no published qualifying conditions
The 2026-07-23 restructure closed one transparency gap and opened another. Frontier capacity is no longer hidden — B200 and B300 now carry public list prices — but the page never says what earns the lower rate: a volume commit, a term length, a reserved-capacity minimum, or simply asking. A buyer reading the table cannot tell whether $1.89/h is realistic for them or a number reserved for eight-figure accounts. Publishing the thresholds — even as bands (“as low as applies above N GPU-hours/month on a 12-month term”) — would make the discount column actionable instead of aspirational, the same discipline good usage-based pricing pages apply to overage rates.
4. The hero headline and the compute table now disagree
The page still leads with “as low as $1.89/hr for H100” directly above a table whose H100 List Price is $3.99/h. Both statements are true, but a buyer scanning the hero forms a price expectation the table then doubles — the kind of mismatch that erodes trust in an otherwise unusually transparent page. Leading with the list price and naming the floor as a discount (“H100 from $3.99/h, as low as $1.89/h on custom deployments”) would carry the same commercial message without the whiplash.
Monetization stack & signals : how Fal builds & buys its revenue engine
Buys 7 Builds 0 11 open roles
fal buys its meter: a dedicated Payments team wires Orb for usage metering and Stripe for payments rather than building either. A widening commercial org (enterprise PM, GTM data, account managers on Salesforce/Gong) overlays sales-led motion onto the self-serve core.
-
“integrate with Orb for usage metering and Stripe for payments and invoicing”
-
“integrate with Orb for usage metering and Stripe for payments and invoicing”
-
“Technologies You'll Use Salesforce • Slack • Hex • Gong • Apollo”
-
“Technologies You'll Use Salesforce • Slack • Hex • Gong • Apollo”
-
“Experience with NetSuite, Carta, Shareworks, or similar tools a plus”
-
“building ingestion pipelines into a warehouse (BigQuery, Snowflake, Redshift) … Strong SQL and working proficiency in dbt”
-
“SQL, Amplitude/Looker, or plain-text CSVs—whatever gets you to the insight fastest”
- Senior Data Scientist, GTM RevOps seen May 13, 2026
- AR Specialist RevOps seen May 13, 2026
- Technical Accounting and Reporting Manager Deal desk seen Mar 19, 2026
- Enterprise Product Manager Monetization seen Dec 15, 2025
- Staff Software Engineer, Payments Billing engineering seen Nov 18, 2025
- Account Manager, Commercial (North America) Customer success seen Oct 17, 2025
- Account Manager, Enterprise Customer success seen Oct 17, 2025
- Staff Software Engineer, Forward Deployed Customer success seen Oct 17, 2025
- Technical Business Development (Model Labs) Customer success seen Oct 17, 2025
- Software Engineer, Growth Growth seen Jul 21, 2025
- Senior Data Scientist, Growth Growth seen Jul 21, 2025
Signals reviewed · derived from public job posts
Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.
Key takeaways
- Price the output, not the seat. fal meters per generated second, image, and megapixel — the exact units its customers produce — so cost scales with usage instead of headcount. Generative-media tools should consider output as the value metric before defaulting to per-user pricing.
- Normalise prices for the buyer. Publishing “output per $1” turns a price list into a comparator and removes the buyer’s math. A small UX choice can materially lower the barrier to purchase.
- Offer both managed and raw surfaces. Selling per-output APIs and per-hour GPU rental side by side captures both convenience-buyers and cost-optimisers without forcing a single model on everyone.
- Publish the band, not just the point. In July 2026 fal replaced one rate per GPU with a List Price and an “As low as” floor, un-gating frontier capacity in the same move. If you need room to negotiate, showing both ends of the spread keeps the page honest and turns “contact sales” into a quantified offer — but only if you also say what earns the lower number.
- Fully variable pricing needs guardrails. Without a free tier, spend caps, or commits, a metered model can produce both trial friction and bill shock; the absence of these controls is the main gap in fal’s otherwise clean model.
UBP implications
- Output-based metering is the natural fit for generative media. fal shows that per-second and per-image billing maps cleanly onto generative AI’s unit of value, a template other media-inference vendors can copy directly.
- Self-published cost normalisation is an emerging UBP best practice. “Output per $1” is a buyer-facing innovation that other usage-based vendors could adopt to make complex per-unit rates legible.
- Pure-usage models trade predictability for fairness — and rate bands are a third option. fal’s seatless, floorless model is maximally fair but maximally variable, so metered pricing usually needs commitment options or spend controls layered on to be enterprise-ready. Its 2026-07 List Price / “As low as” split points at a middle path: publish a range rather than a point, which preserves the metered model while making the negotiated rate visible instead of replacing it with “contact us.”
Sources
- fal pricing page (accessed 2026-08-28)
- fal enterprise page (accessed 2026-08-28)
- fal documentation (accessed 2026-07-23)
- fal compute pricing docs (accessed 2026-07-23)
- Veo 3 model page (accessed 2026-08-28)
- Veo 3 Fast model page (accessed 2026-08-28)
- Veo 3.1 model page (accessed 2026-08-28)
- Wan 2.5 text-to-video model page (accessed 2026-08-28)
- Kling 2.5 Turbo Pro model page (accessed 2026-08-28)
- Seedream V4 model page (accessed 2026-07-23)
- FLUX.1 Kontext [pro] model page (accessed 2026-07-23)
- Nano Banana model page (accessed 2026-08-28)
- Qwen Image model page (accessed 2026-07-23)
- Ovi model page (accessed 2026-07-23)
- fal blog (accessed 2026-07-23)
Bottom line
fal is generative-media inference priced exactly the way its customers produce value: per second of video, per image, per megapixel, and per GPU-hour — with no seats, no subscription, and no free tier. The model-API half of the page doubles as a cost comparator and has not moved in over a year. The compute half changed shape in July 2026: every GPU now carries a List Price and an “As low as” floor, which un-gated Blackwell capacity but roughly doubled what a self-serve buyer sees for an H100 — transparency traded sideways rather than forward.
Want to compare fal against other usage-based AI infrastructure pricing? Browse the pricing blueprint.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
GPU compute split into List Price and 'As low as'
The Serverless & Compute table dropped its per-second column and now publishes two hourly rates per GPU. The rates fal used to advertise as the price became the floor (H100 $1.89/h, H200 $2.10/h) while new list prices sit roughly 2x higher (H100 $3.99/h, H200 $4.50/h), so a self-serve buyer who does not contact sales sees a materially higher number. B200 lost its 'contact us' gate at $6.25/h list ($3.49/h as low as); B300 288GB ($8.50/h, $4.49/h) and RTX PRO 6000 96GB ($2.99/h, $1.10/h) joined the fleet; A100 40GB was retired. Model API rates were unchanged.
Serverless & Compute + Model APIs
The page presents two metered surfaces — Serverless & Compute (GPU per-hour/per-second: H100 $1.89/h, H200 $2.10/h, A100 $0.99/h, B200 'contact us') and Model APIs (Video: Wan 2.5 $0.05/s, Kling 2.5 Turbo Pro $0.07/s, Veo 3 $0.4/s, Ovi $0.2/video; Image: Seedream V4 $0.03, Flux Kontext Pro $0.04, Nanobanana $0.0398, Qwen $0.02/MP). No seats, subscriptions, or free tier.
Modern layout: Output-Based Pricing + B200 'contact us'
Coinciding with the $125M Series C ($1.5B valuation), the page reached its current structure: 'GPU Pricing' (H100 $1.89/h, H200 $2.10/h, A100 $0.99/h, B200 184GB 'contact us') plus 'Output-Based Pricing' split into Video Models and Image Models with an 'Output per $1' comparator column. B200 Blackwell capacity was added as sales-gated.
H100 cut to $1.89/hr; full per-hour fleet
The H100 hourly rate dropped to $1.89/h ($0.0005/s), with H200 141GB at $2.10/h, A100 40GB at $0.99/h ($0.0003/s), and A6000 48GB at $0.60/h. These H100/H200/A100 rates are the same ones live on the page a year later.
Per-hour GPU pricing introduced ($1.99/hr H100)
fal switched its GPU fleet to a per-hour display alongside per-second, leading with 'Get H100s from as low as $1.99/hr.' The new 'GPU Pricing' table priced H100 80GB at $1.99/h. This is the move from a budget-slider comparator to the per-hour/per-second table that defines the current page.
$49M Series B + pricing hero redesign
A site banner announced 'fal Raises $49M Series B to Power the Future of AI Video' (Notable Capital + a16z). The pricing page was redesigned with the 'Fast, reliable, and cost-efficient' hero and simplified nav (Pricing / Enterprise), keeping both the model-output table and the GPU budget comparator (H100 $0.00125/s, A100 $0.00111/s, A6000 $0.000575/s).
H100 added + 'Billing Based on Model Output' table
fal added the H100 80GB ($0.00125/s) to the top of the GPU fleet and introduced a 'Billing Based on Model Output' table (FLUX.1 [dev]/[schnell]/[pro], Stable Diffusion 3 Medium, Stable Video) — 'models below are billed by model output, instead of compute seconds.' The new diamond 'fal' wordmark also debuted. Per-output model-API billing layered on top of GPU rental.
'Choose a budget' output comparator UI
fal added a budget slider ($1–$200) that translated GPU-second rates into 'how many inferences per $20' — e.g. SDXL ~10,296 runs, SDXL Lightning ~47,415 runs, Whisper v3 ~3,677 runs. Per-second GPU rates persisted (A100 $0.00111/s, A6000 $0.000575/s, A10G $0.00053/s) but the page became a cost comparator. Birth of fal's output-normalisation.
Raw per-second unit pricing
fal's earliest captured pricing billed compute by raw unit-second: CPU $0.00003/s, Memory $0.000004/s, GPU A100 $0.001/s, A10G $0.0002/s, T4 $0.00009/s, plus Storage $1/GB/month. A 'Choose your Machine Type' picker summed units (e.g. A100 machine = $0.00111/s). No output-based pricing or budget comparator yet.
- · fal's hero still advertises H100s 'as low as $1.89/hr,' but since July 2026 that rate is the floor of a two-column table whose H100 List Price is $3.99/h — the headline and the price list differ by 2x on the same screen.
- · fal has no seats, no monthly plans, and no free tier on its public pricing page — every line item is metered per output or per unit of compute time.
- · fal normalises model-API prices to 'output per $1' on its own pricing page (e.g. 20 seconds of Wan 2.5 video, or 33 Seedream V4 images), turning the price table into a buyer-facing cost comparator.
Questions & answers
- Does fal.ai have a free tier or monthly plan?
- No. fal's public pricing page lists only usage-based rates — per-output model API calls and per-hour GPU compute. There are no seats, subscriptions, or free tier advertised.
- How much does an H100 GPU cost on fal?
- fal lists the H100 (80GB) at a $4.50/h list price, or as low as $1.89/h. The rest of the fleet: H200 141GB $4.50/h (as low as $2.10/h), B200 180GB $6.25/h (as low as $3.49/h), B300 288GB $8.50/h (as low as $4.49/h), and RTX PRO 6000 96GB $2.99/h (as low as $1.10/h).
- How is fal video generation priced?
- Video models are billed by output unit — per second or per video. The pricing page lists Wan 2.5 at $0.05/second, Kling 2.5 Turbo Pro at $0.07/second, Veo 3 at $0.4/second, and Ovi at $0.2/video. Check the model page before budgeting: Wan 2.5's $0.05/second is the 480p rate ($0.10 at 720p, $0.15 at 1080p). As of 2026-08-13, fal's Veo lineup runs on Veo 3.1 (the old Veo 3 and Veo 3 Fast endpoints are deprecated): Standard is $0.20/second audio off or $0.40/second audio on at 720p/1080p ($0.40/$0.60 at 4K), and Fast is $0.10/second audio off or $0.15/second audio on at 720p/1080p ($0.30/$0.35 at 4K).
- How is fal image generation priced?
- Image models are billed per image or per megapixel. Examples: Seedream V4 at $0.03/image, Flux Kontext Pro at $0.04/image, Nanobanana at $0.0398/image, and Qwen at $0.02/megapixel.
- What does fal enterprise include?
- fal enterprise adds private model hosting, custom inference and training kernels, foundational model research, dedicated serverless infrastructure, SOC2 certification, SSO, and user management — all priced via Contact Sales.