Ask
Deprecation

Groq pulls Llama 4 Scout, Qwen3 32B, and the Browser Automation tool

Groq pricing

Groq removed Llama 4 Scout and Qwen3 32B from its public rate card and docs catalog, dropped the $0.08/hour Browser Automation tool, and swapped enterprise-only Minimax M2.5 for M2.7. Retained model prices held stable.

Before

Eight priced LLMs on groq.com/pricing including Llama 4 Scout (17Bx16E) at $0.11/$0.34 and Qwen3 32B at $0.29/$0.59; five Compound built-in tools including Browser Automation at $0.08/hour; enterprise-only Minimax M2.5 and Qwen3-VL 32B.

After

Six priced LLMs (GPT OSS 20B, GPT OSS Safeguard 20B, GPT OSS 120B, Llama 3.3 70B Versatile, Llama 3.1 8B Instant, Qwen 3.6 27B); four Compound built-in tools (basic search, advanced search, visit website, code execution); one enterprise-only model, Minimax M2.7.

Proof of change

Groq's pricing pages, as we captured them on two dates.

Second source confirmed
Captured Jul 14, 2026 Captured Jul 21, 2026 · 7 days apart
main
$0.34 $0.29
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Show 1 other page we compared
models-docs
$0.34 $0.29
Captured Jul 14, 2026
Captured Jul 21, 2026

Showing the whole page as captured — scroll either panel, or open it at full size.

Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.

Between the 2026-07-14 and 2026-07-21 captures of groq.com/pricing, Groq shrank its published serverless catalog rather than repricing it. Llama 4 Scout (17Bx16E) 128k and Qwen3 32B 131k disappeared from both the pricing page rate card and the Supported Models page in the docs, where they had been listed under Production and Preview models respectively. The enterprise-only column also thinned: Qwen3-VL 32B is gone and Minimax M2.5 was replaced by a newer Minimax M2.7, still at contact-sales.

The agentic tool rate card lost a line too. Browser Automation, which had launched the previous week at $0.08 per hour, is no longer listed under Built-In Tools (Compound); that table is back to basic search at $5 per 1,000 requests, advanced search at $8 per 1,000, visit website at $1 per 1,000, and code execution at $0.18 per hour. The separate GPT-OSS tool table (browser search $5 per 1,000, visit website $1 per 1,000, Python code execution $0.18 per hour) is unchanged.

Every retained price held: Llama 3.1 8B Instant at $0.05/$0.08, GPT OSS 20B and Safeguard 20B at $0.075/$0.30, GPT OSS 120B at $0.15/$0.60, Llama 3.3 70B Versatile at $0.59/$0.79, Qwen 3.6 27B at $0.60/$3.00, Whisper at $0.111 and $0.04 per hour transcribed, and Orpheus TTS at $22.00 and $40.00 per 1M characters. The change is catalog contraction, not a price move — and it is a reminder that Groq’s docs explicitly warn that Preview models “may be discontinued at short notice,” so a model you priced a workload around can leave the rate card inside a week.

From Groq's pricing timeline
Catalog Contraction: Two Models and the Browser Automation Tool Withdrawn

One week after expanding the rate card, Groq contracted it without repricing. Llama 4 Scout (17Bx16E, $0.11/$0.34) and Qwen3 32B ($0.29/$0.59) left both the pricing page and the docs model catalog, the Browser Automation built-in tool ($0.08/hour) was pulled from Built-In Tools (Compound) seven days after launch, Qwen3-VL 32B was dropped from the enterprise-only list, and Minimax M2.5 was replaced by M2.7. Every retained price held exactly, so the cost impact lands as forced substitution rather than as a rate move — the nearest remaining Qwen, Qwen 3.6 27B at $0.60/$3.00, costs 2.1x more on input and 5.1x more on output than the delisted Qwen3 32B.

About Groq
groq.com ↗

Groq runs a pure-usage per-token serverless inference API on its proprietary LPU silicon. As of 2026-08-11 the dedicated marketing pricing page (groq.com/pricing/) has been removed and redirects to the homepage; rates survive only in the GroqCloud developer docs. As of 2026-08-26, Llama 3.1 8B Instant and Llama 3.3 70B Versatile — previously $0.05/$0.08 and $0.59/$0.79 per 1M tokens — moved to Enterprise-only "Contact Sales" pricing and dropped off both the Free and Developer rate-limit tables. The self-serve catalog is now GPT OSS 20B and Safety GPT OSS 20B (formerly "GPT OSS Safeguard 20B") at $0.075/$0.30 (1,000 T/SEC), GPT OSS 120B at $0.15/$0.60 (500 T/SEC), Qwen 3.6 27B at $0.60/$3.00, and two Preview moderation models — Llama Prompt Guard 2 22M and Prompt Guard 2 86M — at $0.03/$0.03 and $0.04/$0.04 per 1M tokens. As of 2026-08-27, Qwen 3.8-27B joined the Preview catalog with a published price of $0.80/$4.00 per 1M tokens (450 T/SEC) — the only change versus the prior day's capture.

Free tier
Yes
Commits
Available
Transparency
public

Groq pricing history

  1. Aug 2026
    Qwen 3.8-27B Gains a Published Preview-Tier Price
  2. Aug 2026
    Llama 3.1 8B Instant and Llama 3.3 70B Versatile Moved to Enterprise-Only Pricing
  3. Aug 2026
    Public Pricing Page Removed — Rates Survive Only in Developer Docs
  4. Jul 2026
    Catalog Contraction: Two Models and the Browser Automation Tool Withdrawn
  5. Jul 2026
    Text-to-Speech SKU + Browser Automation Tool + New Models
Full Groq timeline

More Groq activity

All pricing activity