Ask
Launch

fal launches Veo 3.1 and cuts Veo-class video pricing ~50-60%

Fal pricing

fal added Google's Veo 3.1 (4K, native audio, video extension) and repriced the whole Veo line: Standard drops to $0.20/s no-audio ($0.40/s with audio) from $0.50/$0.75, Fast to $0.10/s ($0.15/s) from $0.25/$0.40. The old Veo 3 and Veo 3 Fast endpoints are now deprecated.

Before

Veo 3 (fal-ai/veo3): $0.50/s audio off, $0.75/s audio on. Veo 3 Fast (fal-ai/veo3/fast): $0.25/s audio off, $0.40/s audio on. Only 1080p, 8s clips, no image-to-video or extend-video modes.

After

Veo 3.1 (fal-ai/veo3.1) Standard: $0.20/s audio off, $0.40/s audio on at 720p/1080p ($0.40/$0.60 at 4K). Veo 3.1 Fast: $0.10/s audio off, $0.15/s audio on at 720p/1080p ($0.30/$0.35 at 4K). Adds true 4K output, image-to-video, first/last-frame, reference-to-video, and extend-video (up to ~148s). The old fal-ai/veo3 and fal-ai/veo3/fast endpoints now show 'This endpoint is deprecated — This model is no longer supported' and display the same lower Veo 3.1-equivalent rates rather than their original $0.50/$0.75 and $0.25/$0.40 prices.

fal added Google DeepMind’s Veo 3.1 across nine endpoints (text-to-video, image-to-video, first/last-frame, reference-to-video, and extend-video, each in Standard and Fast tiers) and retired the original Veo 3 and Veo 3 Fast endpoints in the same move. The old endpoints still resolve, but each now carries a “This endpoint is deprecated — This model is no longer supported” banner and quotes the new, lower Veo 3.1-equivalent per-second rate rather than the price its own marketing copy still describes.

The repricing is steep: Standard-tier video without audio falls from $0.50/s to $0.20/s (a 60% cut), and with audio from $0.75/s to $0.40/s (47%) at 720p/1080p — fal’s new 4K tier costs $0.40/s and $0.60/s respectively, roughly what the old 1080p-only Standard rate used to be. Fast-tier drops similarly, from $0.25/$0.40 per second to $0.10/$0.15 at 720p/1080p. Veo 3.1 also adds capabilities the old endpoint never had: true 4K output, image-to-video and reference-to-video modes, and an extend-video endpoint that chains clips up to roughly 148 seconds from a single 8-second starting generation.

Notably, fal’s top-level /pricing summary table was not touched by this change — it still shows a single “Veo 3” row at $0.4/second, unchanged since 2025-07. A buyer who never clicks into the model page would not see either the new Veo 3.1 launch or the underlying price cut; the detail lives entirely on fal’s per-model pages.

From Fal's pricing timeline
GPU compute split into List Price and 'As low as'

The Serverless & Compute table dropped its per-second column and now publishes two hourly rates per GPU. The rates fal used to advertise as the price became the floor (H100 $1.89/h, H200 $2.10/h) while new list prices sit roughly 2x higher (H100 $3.99/h, H200 $4.50/h), so a self-serve buyer who does not contact sales sees a materially higher number. B200 lost its 'contact us' gate at $6.25/h list ($3.49/h as low as); B300 288GB ($8.50/h, $4.49/h) and RTX PRO 6000 96GB ($2.99/h, $1.10/h) joined the fleet; A100 40GB was retired. Model API rates were unchanged.

About Fal
fal.ai ↗

fal (fal.ai) is a generative-media inference platform that prices purely on usage: serverless per-output model APIs plus dedicated GPU compute, with no seats, no subscriptions, and no free tier on its public pricing page.

Pricing model pure usage
Sales motion self servesales led
Free tier
No
Commits
None
Transparency
public

Fal pricing history

  1. Jul 2026
    GPU compute split into List Price and 'As low as'
  2. Jun 2026
    Serverless & Compute + Model APIs
  3. Jul 2025
    Modern layout: Output-Based Pricing + B200 'contact us'
  4. May 2025
    H100 cut to $1.89/hr; full per-hour fleet
  5. Apr 2025
    Per-hour GPU pricing introduced ($1.99/hr H100)
Full Fal timeline

More Fal activity

All pricing activity