Groq moves Llama 3.1 8B and Llama 3.3 70B to Enterprise-only pricing
Groq's two Llama models lost their public per-token price and now show Contact Sales in the docs, dropping off the Free and Developer rate-limit tables entirely.
Llama 3.1 8B Instant self-serve at $0.05 input / $0.08 output per 1M tokens (560 T/SEC); Llama 3.3 70B Versatile self-serve at $0.59 input / $0.79 output per 1M tokens (280 T/SEC); both listed with public Free and Developer plan rate limits.
Llama 3.1 8B Instant and Llama 3.3 70B Versatile marked "Enterprise" with "Contact Sales" in place of both price and rate limits; both models absent from the Free Plan Limits and Developer Plan Limits tables. GPT OSS 20B ($0.075/$0.30) is now the cheapest self-serve model.
Groq’s GroqCloud developer docs (console.groq.com/docs/models) gated its two Llama models behind an Enterprise sales conversation on this capture, the most consequential change since the dedicated groq.com/pricing/ page was retired on 2026-08-11. Llama 3.1 8B Instant and Llama 3.3 70B Versatile — previously the platform’s headline low-cost, high-throughput models — now show “Contact Sales” in both the price and rate-limit columns, and a cross-check against the docs’ Rate Limits page confirmed neither model appears on the Free Plan Limits or Developer Plan Limits tables anymore. Every other production price (GPT OSS 120B, GPT OSS 20B, Whisper, Whisper Turbo) held unchanged. The same capture added two small Preview-tier moderation models (Llama Prompt Guard 2 22M and Prompt Guard 2 86M, both priced at $0.03–$0.04 per 1M tokens) and formalized “Groq Compound” and “Compound Mini” as named Production Systems with no published per-token price.
The two Llama models on Groq's rate card — previously self-serve at $0.05/$0.08 and $0.59/$0.79 per 1M tokens — are now marked "Enterprise" in the GroqCloud docs model catalog, with "Contact Sales" printed in place of both the price and rate-limit columns. Confirmed by a second surface: both models are also absent from the Free Plan Limits and Developer Plan Limits tables on the docs' Rate Limits page, where they previously had public RPM/TPM caps. GPT OSS 20B ($0.075/$0.30) is now the cheapest self-serve model on the card. Every other production price (GPT OSS 120B, Whisper, Orpheus, Qwen 3.6 27B) held unchanged. The docs catalog also grew two new Preview moderation models (Llama Prompt Guard 2 22M at $0.03/$0.03, Prompt Guard 2 86M at $0.04/$0.04 per 1M tokens) and formalized "Groq Compound" / "Compound Mini" as named Production Systems with no published per-token price.