AI Summary
About
Cerebras Systems is a Santa Clara-based AI hardware and cloud inference company founded in 2016 by Andrew Feldman (CEO) and Gary Lauterbach (CTO). The company’s central innovation is the Wafer Scale Engine (WSE): a single silicon die occupying an entire semiconductor wafer, containing up to 4 trillion transistors and 900,000 AI-optimized cores. By eliminating the inter-chip communication bottleneck that limits GPU clusters, the WSE achieves dramatically higher throughput for large model inference at lower latency.
Cerebras operates two distinct product lines. The first is Cerebras Inference: a public cloud API that delivers hosted LLM inference on open-source models (GPT-OSS, Gemma, GLM, and — via Dedicated Endpoints — Llama, Qwen, Mistral, DeepSeek and more) at speeds the company markets as 20× faster than OpenAI and Anthropic, billed per million input and output tokens and entered through a $5 free-credit trial. A third, narrower SKU sits alongside it: Cerebras Code, a fixed-price monthly coding subscription with a daily token allowance. The second product line is the CS-3 compute system: a rack-scale appliance housing the WSE-3 chip, sold under enterprise and government contracts for on-premises or managed deployment by research institutions, national laboratories, healthcare systems, and large enterprises.
On August 18, 2026, Cerebras announced CS-4 — a fourth-generation system built from three Wafer Scale Engine 3 Turbo processors on a redesigned modular rack architecture (“Nexus Platform”), claimed up to 30x faster than GPU systems and up to 2x faster than CS-3, with first shipments beginning in Q3 2026. Like CS-3, CS-4 pricing is not publicly disclosed and access runs through direct enterprise sales.
The company raised approximately $750M in total venture funding prior to its 2024 IPO attempt. Cerebras filed its S-1 in August 2024 targeting an ~$8 billion valuation, but the IPO was withdrawn in November 2024 after CFIUS opened a national-security review related to G42, the UAE-based AI conglomerate that had been Cerebras’s largest customer and which had prior ties to Huawei. Cerebras subsequently raised a private funding round at a comparable valuation to continue operations while the regulatory situation resolved. By 2026 the company reports annual recurring revenue in the $100M–$500M range, driven primarily by enterprise hardware contracts and growing inference API revenue.
Pricing summary : Tiered API access, per-token rates, and fixed coding plans
Cerebras Inference is sold through three access tiers plus a separate coding subscription. The entry point is a Free Trial that grants $5 in free credits after making an account, with access to all Cerebras powered models and community support via Discord — a one-time credit grant rather than an open-ended free allowance. The docs attach two conditions the marketing card omits: the credits arrive only “after adding a verified payment method” (there is no charge for adding one), and they “expire 30 days after they’re granted”. The self-serve Developer tier starts at just $10 and unlocks 10x higher rate limits and higher-priority processing; the docs describe the transition more simply — “your first purchase moves you to the Developer tier”, which also removes the free trial’s hourly and daily token caps. The Enterprise tier (contact sales) adds the highest rate limits, lowest latency via dedicated queue priority, support for custom model weights, model fine-tuning and training services, and a dedicated support team with response-time guarantees. Underlying consumption is billed per million input and output tokens against the published rate card.
The public “Developer Tier Pricing” rate card lists two models as of 2026-08-26, down from three a month earlier — ZAI GLM 4.7 was removed on its previously footnoted deprecation date and no longer appears in either the pricing page’s rate table or the docs Model Catalog:
- GPT OSS 120B — $0.35 input / $0.75 output, ~3000 tokens/s. The only model the docs classify as a Production model.
- Google Deepmind Gemma 4 31B — $0.99 input / $1.49 output, ~1,800 tokens/s (~1850 in the docs). Classified Preview.
The page footnotes that “Preview models are intended for evaluation purposes only, and are not intended for use in production environments” — so of the two public rate-card models, only GPT OSS 120B is sanctioned for production. A much wider catalog (Qwen3 and Qwen3-Coder, Llama 3 and Llama 4, Mistral Small/Large 3/Devstral 2/Mixtral, DeepSeek V3.X, Kimi K2.X, MiniMax M2.X, GLM 4.X and 5.X (now including the former public-card model GLM 4.7), Gemma 4, StepFun Step 3.X Flash, ByteDance OSS Seed, ServiceNow Apriel) is reachable only through Dedicated Endpoints on reserved-capacity custom pricing.
Separately, Cerebras Code is a fixed-price coding subscription, now sold entirely from its own cerebras.ai/code product page rather than the main pricing page (which no longer lists Code plans at all as of the 2026-08-26 capture, a change from May/July 2026 when Code cards appeared directly on /pricing): Pro at $50/month (send up to 24 million tokens/day, “enough for 3–4 hours of uninterrupted vibe coding”) and Max at $200/month (up to 120 million tokens/day, for heavy coding workflows). Both tiers display SOLD OUT, as does a third Free tier at $0 (“GLM 4.7 access with limited tokens and requests”) that only the product page lists. That page names GLM 4.7 as the model behind all three Code plans, which is the model whose separate public per-token rate-card entry was removed on its Aug 17, 2026 deprecation date. The CS-3 hardware product remains an entirely separate commercial motion: enterprise contracts negotiated by a direct sales team, with per-unit pricing not publicly disclosed. On 2026-08-18 Cerebras announced CS-4, a fourth-generation system, under the same undisclosed-pricing, contact-sales motion. Cerebras also resells API access through partner channels — AWS Marketplace, OpenRouter, Hugging Face, and Vercel.
What makes this different: Cerebras charges for speed — but doesn’t charge a speed premium. Every token processed through Cerebras Inference arrives 10–20× faster than GPU-cloud equivalents at pricing that matches or undercuts those slower alternatives. This inversion of the traditional cost/performance tradeoff is the core commercial proposition. See how AI inference providers are restructuring their pricing models for context on why this matters.
Pricing by product
Cerebras Inference — access tiers
| Tier | Price | What you get | Sales motion |
|---|---|---|---|
| Free Trial | $5 in free credits (one-time, after creating an account) | Access to all Cerebras powered models; community support via Discord. Docs caveats: credits are granted only “after adding a verified payment method” (no charge for adding one) and “expire 30 days after they’re granted”; capped at 5 requests/min, 30K tokens/min and 1M tokens/day | Self-serve |
| Developer | From $10 (self-serve) | Everything in Free, plus 10x higher rate limits than free tier and higher priority processing. Docs: “your first purchase moves you to the Developer tier … and no hourly or daily token caps” (no minimum purchase amount is stated in the docs) | Self-serve / PLG |
| Enterprise | Contact sales | Everything in Developer, plus highest rate limits for production workloads, lowest latency with dedicated queue priority, support for custom model weights, model fine-tuning and training services, dedicated support team with response-time guarantees | Sales-led |
Cerebras Inference — public “Developer Tier Pricing” rate card
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Listed speed | Status / caveat |
|---|---|---|---|---|
| GPT OSS 120B | $0.35 | $0.75 | ~3000 tokens/s | Production (120B params) |
| Google Deepmind Gemma 4 31B | $0.99 | $1.49 | ~1,800 tokens/s | Preview — evaluation only, not for production (31B params) |
| — | — | — | Removed 2026-08-17 — deprecated as previously scheduled (see Pricing evolution for its former rate); no longer on the public rate card or docs Model Catalog as of 2026-08-26. Still reachable via Dedicated Endpoints under Z.AI GLM 4.X/5.X |
Footnoted on the pricing page: “Preview models are intended for evaluation purposes only, and are not intended for use in production environments.” The docs Model Catalog now lists only two models — GPT OSS 120B (Production) and Gemma 4 31B (Preview) — confirmed via a fresh capture of both the pricing page rate table and the docs Model Catalog sidebar/table on 2026-08-26. Public-endpoint models are available on the free trial and pay-as-you-go tiers, subject to rate limits.
Cerebras Inference — Dedicated Endpoints (reserved capacity)
| Item | Detail |
|---|---|
| Price | Custom — “Dedicated endpoints are available to enterprise customers. Contact us.” |
| Model families | Alibaba Qwen (Qwen3, Qwen3-Coder), OpenAI GPT-OSS, MiniMax M2.X, Google Gemma 4, Meta Llama 3 and Llama 4, Mistral (Small, Large 3, Devstral 2, Mixtral), Z.AI GLM 4.X and GLM 5.X, Moonshot AI Kimi K2.X, DeepSeek V3.X, StepFun Step 3.X Flash, ByteDance OSS Seed, ServiceNow Apriel; multimodal “coming soon” |
| Included features | Fine-tuning (bring your own weights), Management API, Batch API, service tiers for request prioritization, Prometheus-compatible metrics |
| Why buy | Reserved private capacity so latency/throughput are unaffected by other users |
Cerebras Code — coding subscriptions
| Plan | Price | Daily token allowance | Target | Availability |
|---|---|---|---|---|
| Free | $0 | ”GLM 4.7 access with limited tokens and requests” | Trying out Cerebras inference or building a small demo in an AI code editor | SOLD OUT |
| Pro | $50/month | Send up to 24 million tokens/day (“enough for 3–4 hours of uninterrupted vibe coding”) | Indie devs, simple agentic workflows, weekend projects | SOLD OUT |
| Max | $200/month | Send up to 120 million tokens/day | Full-time development, IDE integrations, code refactoring, multi-agent systems | SOLD OUT |
All three Cerebras Code plans, including the Free tier, now live only on the Cerebras Code product page — the main cerebras.ai/pricing page no longer lists any Cerebras Code plans as of the 2026-08-26 capture (a change from the May/July 2026 captures, when Pro and Max cards appeared directly on /pricing). That product page still names GLM 4.7 as the model behind all three Code plans as of the 2026-08-26 capture — even though GLM 4.7’s entry on the public per-token “Developer Tier Pricing” rate card was removed on its previously disclosed Aug 17, 2026 deprecation date. Cerebras Code appears to route to GLM 4.7 independently of the metered rate card (Code is a fixed-price subscription, not per-token billing). Note: the Code page’s small top banner still reads “NOW UPGRADED WITH GLM 4.6,” inconsistent with the page body’s repeated references to GLM 4.7 — an apparent unresolved copy inconsistency on Cerebras’s own page, not something this page is asserting as fact.
Cerebras CS-3 / CS-4 Compute Systems (enterprise hardware)
| Product | Use case | Pricing model | Notes |
|---|---|---|---|
| CS-3 on-premises | In-house training + inference | Enterprise contract | Multi-year; pricing not public |
| CS-3 managed cloud | Cloud-connected dedicated cluster | Enterprise contract | Includes managed ops |
| Cerebras Model Studio | Hosted enterprise inference on private models | Negotiated usage | Custom deployment on CS-3 |
| CS-4 (announced 2026-08-18) | 4th-gen wafer-scale system; 3x WSE-3 Turbo processors on a new modular “Nexus Platform” rack | Enterprise contract, “contact us” | No public price disclosed; claimed up to 30x faster than GPU systems, up to 2x faster than CS-3; first shipments Q3 2026 |
Partner resale channels
| Channel | How it is sold |
|---|---|
| AWS Marketplace | ”Buy with AWS” — test workloads and move to production with “simple controls and flexible pricing” |
| OpenRouter | Cerebras inference reachable through OpenRouter’s unified API |
| Hugging Face | Cerebras-powered models called directly from the Hugging Face Hub |
| Vercel | Deploy and scale Cerebras inference endpoints on Vercel |
Sales motions across products: PLG / self-serve for the Inference API ($5 trial credits → $10 Developer tier) and Cerebras Code subscriptions; sales-led for Dedicated Endpoints, CS-3 hardware systems, and enterprise managed deployments; partner-led resale via AWS Marketplace, OpenRouter, Hugging Face, and Vercel.
Hidden costs : What developers actually pay beyond the base rate
Archetype A: Startup building a customer-facing chatbot on GPT-OSS-120B
A startup routing 10 million input tokens and 3 million output tokens per day through Cerebras for a customer-facing support chatbot, on the public per-token rate card (GPT-OSS-120B, $0.35 in / $0.75 out):
| Line item | Monthly cost |
|---|---|
| Input tokens: 300M × $0.35/1M | approximately $105.00 |
| Output tokens: 90M × $0.75/1M | approximately $67.50 |
| Rate-limit overages / retry overhead (~5%) | approximately $9.00 |
| Estimated total | approximately $181/month |
A team writing or iterating on code rather than serving chat traffic might instead reach for a Cerebras Code Pro subscription at $50/month (up to 24M tokens/day), which converts heavy daily coding usage into a fixed, predictable bill rather than metered per-token spend — though as of writing both Cerebras Code tiers were sold out, pushing those users back to the metered Developer tier.
Archetype B: Research team running long-context batch processing on GPT-OSS-120B
A research team running nightly batch jobs: 1 billion input tokens, 200 million output tokens per month, using GPT-OSS-120B for structured reasoning tasks:
| Line item | Monthly cost |
|---|---|
| Input tokens: 1B × $0.35/1M | $350.00 |
| Output tokens: 200M × $0.75/1M | $150.00 |
| Context overhead (system prompts per call, ~10%) | ~$50.00 |
| Estimated total | ~$550/month |
Note: the docs put GPT-OSS-120B at a 131K context window and 40K max output on paid tiers (65K context / 32K max output on the free trial). Teams generating outputs longer than the max-output ceiling will need to chain calls, increasing input token costs on subsequent turns by feeding prior output as context.
Use the Cerebras pricing calculator to model your own monthly cost based on model selection, token volume, and input/output ratio.
Pricing evolution : From wafer-scale hardware vendor to inference API competitor
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2024 Q3 | 0 | 2 | Cerebras Inference public beta launched; Llama 3.1 8B and 70B added at initial rate card |
| 2024 Q4 | 0 | 0 | IPO blocked by CFIUS; inference API remained stable; Llama 3.3 70B pricing not yet published |
| 2025 Q1 | 1 | 1 | Llama 3.3 70B added at $0.85/$1.20 (premium over Llama 3.1 70B at $0.60/$0.60); first asymmetric input/output pricing |
| 2025 Q2 | 0 | 2 | GPT-OSS-120B ($0.35/$0.75) and Qwen-3-32B ($0.40/$0.80) added |
| 2025 Q3 | 0 | 2 | ZAI-GLM-4.6 ($2.25/$2.75) and ZAI-GLM-4.7 ($2.25/$2.75) added — first premium-priced models |
| 2026 Q1 | 0 | 0 | ZAI-GLM-4.6 deprecated January 2026; ZAI-GLM-4.7 remains active |
| 2026 Q2 | — | 2 | Cerebras Code Pro ($50/mo) and Max ($200/mo) coding subscriptions launched; public per-token rate card narrowed to GPT-OSS-120B + ZAI-GLM-4.7, with Llama/Qwen3 moved to Dedicated Endpoints |
| 2026 Q3 | 0 | 2 | Free tier replaced by a $5-credit Free Trial (packaging, not rate, change); Gemma 4 31B added at $0.99/$1.49 as a Preview model (July); ZAI-GLM-4.7 removed from the public rate card on its disclosed Aug 17, 2026 deprecation date (confirmed via capture on Aug 26); CS-4 hardware system announced Aug 18, 2026 (no public pricing) |
Tracked range: 2024 Q3–2026 Q3. Per-token rates and tier details sourced from cerebras.ai/pricing and inference-docs.cerebras.ai. Quarters not listed above were verified stable.
Notable changes
- 2024-08-29 — Cerebras Inference launched in public beta with Llama 3.1 8B ($0.10/$0.10) and Llama 3.1 70B ($0.60/$0.60). Both models offered at symmetrical input/output pricing, a simplification common in early-stage inference APIs.
- 2024-11-26 — IPO withdrawal announced following CFIUS review. No pricing changes accompanied the event; the company continued to operate Cerebras Inference unchanged.
- 2025 Q1 — Llama 3.3 70B introduced at $0.85 input / $1.20 output — the first asymmetric price pair in the Cerebras catalog, reflecting the standard industry shift toward differentially priced output tokens as generation is more compute-intensive than prefill.
- 2025-05 — GPT-OSS-120B added at $0.35/$0.75; notably cheaper input pricing than Llama 3.3 70B despite being a larger model, reflecting Cerebras’s strategic interest in establishing itself as the fastest platform for OpenAI’s open-weight model.
- 2025 Q3 — ZAI-GLM-4.6 and 4.7 (Zhipu AI multilingual models) added at $2.25/$2.75 — a 2.6–2.3× premium over Llama 3.3 70B, representing the first specialized-model premium on the platform.
- 2026-01-20 — ZAI-GLM-4.6 deprecated; ZAI-GLM-4.7 continues as the sole ZAI model.
- 2026 Q2 — Cerebras restructured its commercial model: three access tiers (Free, a self-serve Developer tier from $10, and Enterprise), a public per-token rate card narrowed to GPT-OSS-120B ($0.35/$0.75) and the Preview-only ZAI-GLM-4.7 ($2.25/$2.75), Llama and Qwen3 families relocated to Dedicated Endpoints on custom pricing, and two new fixed-price Cerebras Code coding subscriptions — Pro at $50/month and Max at $200/month — both of which sold out at launch.
- 2026-07-21 — The free entry point stopped being a tier and became a trial. The card previously read “Free — the easiest way to get started with Cerebras,” an open, rate-limited allowance across every public model; it now reads “Free Trial — get started with $5 in free credits after making an account.” The docs are stricter than the card: credits are granted only “after adding a verified payment method” and “expire 30 days after they’re granted”, so the trial is time-boxed as well as balance-boxed. No per-token rate moved, but the economics of evaluation changed completely: $5 buys roughly 14M input or 6.7M output tokens on GPT-OSS-120B, after which there is no free path back onto the platform. The free-to-paid gate shifted from a soft one (rate limits, which throttle but never stop you) to a hard one (credit exhaustion).
- 2026-07-21 — Google Deepmind Gemma 4 31B was added to the public rate card at $0.99 input / $1.49 output per million tokens (~1,800 tokens/s), slotting between GPT-OSS-120B and ZAI-GLM-4.7 on price. It arrived classified Preview — evaluation only — so the card grew from two models to three without adding a second production-sanctioned option.
- 2026-07-21 — ZAI-GLM-4.7 held its $2.25/$2.75 rates but picked up a footnoted deprecation date of Aug 17, 2026, with a migration guide in the docs. Read together with the Gemma addition, the public card now carries one production model and two Preview models, one of which has a published expiry — the production surface of the rate card is a single SKU, and Cerebras is signalling deprecations further ahead than it did when ZAI-GLM-4.6 disappeared in January 2026.
- 2026-08-17 — ZAI-GLM-4.7 was removed from the public “Developer Tier Pricing” rate card and the docs Model Catalog, exactly as its deprecation footnote had disclosed a month earlier. A capture taken 2026-08-26 confirms both the pricing page table and the docs Model Catalog sidebar/table are down to two models (GPT-OSS-120B, Gemma 4 31B); the pricing page’s scroll-table arrows are disabled, confirming there is no additional row hidden behind pagination. GLM 4.7 remains reachable only via Dedicated Endpoints (Z.AI GLM 4.X/5.X families) on custom pricing, and continues to power the fixed-price Cerebras Code subscriptions independent of the metered rate card.
- 2026-08-18 — Cerebras announced CS-4, its fourth-generation compute system, built from three Wafer Scale Engine 3 Turbo processors on a new modular “Nexus Platform” rack architecture. Claimed up to 30x faster than GPU systems and up to 2x faster than CS-3, with up to 10x more throughput per watt. No public pricing was disclosed; first shipments were described as beginning “this quarter” (Q3 2026). Access continues through the same enterprise contact-sales motion as CS-3.
What’s unique : Speed-first pricing that inverts the inference cost/performance curve
1. The price-speed inversion: faster costs less than slower. On GPU-based inference platforms (Together AI, Fireworks AI, Replicate), higher throughput typically requires reserved capacity or higher pricing tiers. Cerebras Inference delivers 10–20× faster token generation at the same or lower per-token price as GPU alternatives. This is not a promotional rate — it reflects the WSE’s architectural efficiency advantage: on-chip SRAM eliminates the HBM memory bandwidth bottleneck that forces GPU inference to batch tokens slowly. For developers building interactive applications where latency is a product feature, Cerebras offers a genuinely different cost/performance profile. See choosing the right usage metric for how latency shapes value metric selection.
2. Symmetric vs. asymmetric pricing as a maturity signal. Early Cerebras models (Llama 3.1 8B, Llama 3.1 70B) carried identical input and output token prices — a simplification that underprices output generation relative to actual compute cost. Newer models (Llama 3.3 70B, GPT-OSS-120B, Qwen-3-32B) have adopted the industry-standard asymmetric structure where output costs 1.5–3× more than input. This evolution mirrors the broader shift in AI pricing models as inference operators get a better handle on their actual per-token compute cost curves.
3. The demo is still ungated by model, but it is now metered by credit. Cerebras’s entry point remains model-complete — every public model is reachable from the free entry, so a developer experiences GPT-OSS-120B at ~3,000 tokens/second rather than a degraded demo. What changed on 2026-07-21 is the meter behind it: the open, rate-limited Free tier became a Free Trial worth $5 in one-time credits. The demonstration logic survives (speed sells itself in the first few minutes), but the free tier stopped being a place anyone could live. That is a deliberate narrowing: at ~14M input tokens of GPT-OSS-120B, $5 is enough to prove latency in a real application and not enough to run a hobby project indefinitely. Cerebras is now converting on credit exhaustion rather than on sustained friction. This aligns with PLG strategies in AI infrastructure.
4. Vertical integration from chip to API, no third-party silicon dependency. Unlike every other LLM inference API (OpenAI, Anthropic, Groq, Together AI — all running on Nvidia GPUs or TPUs), Cerebras controls the full stack from silicon to API. This vertical integration provides pricing stability independent of Nvidia supply chain constraints and licensing costs. It also enables custom optimization at the hardware-software interface that no GPU-based operator can replicate. For enterprise buyers evaluating infrastructure lock-in risk, this is a meaningful architectural differentiator.
5. Dual product line creating two distinct commercial motions. Cerebras operates two fundamentally different businesses under one brand: a consumption-API business (Cerebras Inference) targeting developers with PLG acquisition, and a hardware enterprise business (CS-3 systems) targeting research institutions with multi-year sales cycles. This dual structure creates distinct revenue streams — recurring inference API revenue plus lumpy hardware contract revenue — a combination that complicates financial modeling but reduces customer concentration risk over time.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Fastest publicly available LLM inference by a significant margin (10–20× vs. GPU clouds) | Narrow model catalog limited to open-source models; no access to proprietary models (GPT-4o, Claude, Gemini) |
| Competitive per-token pricing — matches or undercuts GPU-based competitors at equivalent quality | G42 / CFIUS situation created revenue concentration risk and blocked a liquidity event for investors |
| Entry trial exposes all models at full speed, so the core claim can be evaluated on real workloads rather than a degraded demo | Hardware (CS-3) pricing is opaque — enterprises cannot self-serve a cost estimate |
| No Nvidia dependency — pricing independent of GPU supply chain fluctuations | The perpetual free tier ended on 2026-07-21: $5 of one-time trial credit (~14M GPT-OSS-120B input tokens) is now the whole free allowance, it requires a verified payment method up front, and it expires 30 days after grant — so prototypes hit a hard wall rather than a throttle |
| Vertical integration from silicon to API enables proprietary speed optimizations | Only one of the three public rate-card models (GPT OSS 120B) is production-sanctioned; Gemma 4 31B is Preview and ZAI-GLM-4.7 is dated for deprecation on Aug 17, 2026 |
| OpenAI-compatible API allows drop-in replacement for applications already using OpenAI SDK | Limited enterprise cost controls in the API layer: the console documents usage, cost and audit logs, but no spend caps, budgets, or spend alerts |
Billing UX : Self-serve API keys with token-level usage metering
- Get API key — Both the Free Trial and Developer cards route to a self-serve GET API KEY button; only the Enterprise card routes to CONTACT SALES. The trial issues $5 in free credits after making an account — the docs add that the credits land only “after adding a verified payment method” (“there is no charge for adding a payment method — you only pay if you choose to purchase additional credits”) and that they “expire 30 days after they’re granted”.
- Cloud Console — The documented billing surface is the Cerebras Cloud Console, with dedicated sections for Projects, API Keys, Playground, Usage & Monitoring, and Account & Billing. Credits are bought through the Billing tab, and “your first purchase moves you to the Developer tier”.
- Usage analytics — The console’s analytics dashboard carries three tabs — Usage (requests and token consumption over a date range), Cached-Usage (cache hit rate and token breakdown), and Cost (monthly spend by model, delayed up to 10 minutes) — plus request logs, admin-only audit logs of console actions, and a Limits page showing per-model per-minute and per-day ceilings. There is still no documented spend cap, budget, or spend-alert control.
- Rate limits — Rate limits are a first-class documented control (a Rate Limits page sits under Get Started in the docs). Public-endpoint models are “available on the free trial and pay-as-you-go tiers, subject to rate limits”, and the Developer tier’s headline benefit is “10x higher rate limits than free tier”.
- API compatibility — Cerebras Inference exposes an OpenAI-compatible API (documented under Compatibility → OpenAI Compatibility, plus Python and Node.js SDKs), so switching providers is a base-URL and key change.
- Cost-control capabilities — The docs list Prompt Caching, Predicted Outputs (Preview), and Payload Optimization as capability pages; these are the levers that reduce billable token volume on the shared endpoints.
- Dedicated Endpoints controls — Reserved-capacity customers additionally get a Management API (programmatically manage models, capacity, and endpoints), a Batch API for asynchronous bulk workloads, configurable Service tiers for request prioritization, and Prometheus-compatible metrics for requests, tokens, latency, and health.
- Deprecation policy surface — The docs carry explicit Deprecations, Change Log, Preview Releases, Service Status, and Error Codes pages. The previously banner-flagged ZAI GLM 4.7 deprecation took effect on its disclosed Aug 17, 2026 date — the model no longer appears in the docs Model Catalog or its sidebar as of the 2026-08-26 capture.
- Partner billing — API access can alternatively be purchased and billed through AWS Marketplace, OpenRouter, Hugging Face, or Vercel, moving the invoice relationship to the partner.
- Enterprise CS-3 billing — Hardware contracts are managed through a separate enterprise relationship (order-to-cash runs on NetSuite per hiring signals), not through the Cloud Console.
Strategic wins : Why Cerebras’s pricing decisions have worked
1. Positioning speed as the value metric, not capability
When Cerebras launched Cerebras Inference in August 2024, the inference API market was already crowded with providers offering access to the same Llama models. Rather than competing on model selection (where every GPU-cloud provider had the same catalog) or price (where margins are thin across the board), Cerebras competed on a dimension its hardware uniquely owned: speed. By demonstrating 2,100 tokens/second on Llama 3.1 8B — more than 20× faster than GPU alternatives — Cerebras created a product moment that no competitor could immediately replicate. The decision to charge at parity with slower competitors reinforced the narrative: “same price, radically faster.” This value-metric differentiation generated significant developer attention and organic distribution at launch.
2. An unmetered free tier bought the audience — then Cerebras closed it once it had one
For its first two years Cerebras exposed all public models on an open, rate-limited free tier, letting any developer feel the speed advantage in minutes. The rate limiting was strict enough to push production workloads to paid but loose enough for genuine latency evaluation, and that PLG-first onboarding drove organic adoption through developer communities.
The win is worth reading as a completed play rather than a standing policy. On 2026-07-21 the free tier became a $5-credit Free Trial, which is what a company does once acquisition is no longer the binding constraint and the free traffic is consuming scarce wafer-scale capacity — the same capacity pressure visible in both Cerebras Code tiers sitting SOLD OUT. The strategic lesson holds either way: an unmetered free tier is a cheap way to buy attention in a commoditized market, and it is a cost that a hardware-constrained provider eventually has to retire. The risk now is that the goodwill was priced into the developer relationship, and a trial that expires converts differently than a tier that never did.
3. OpenAI API compatibility eliminated switching cost
By building Cerebras Inference as an OpenAI-compatible API, the company ensured that any developer already using the OpenAI SDK, LangChain, LlamaIndex, or other OpenAI-compatible frameworks could switch to Cerebras by changing exactly two lines of code: the base URL and the API key. This technical decision eliminated the switching cost objection entirely — a developer who wants to test Cerebras can do so against their existing application without any refactoring. The strategy mirrors how Groq and Together AI grew initial developer adoption and is now table stakes for new inference providers.
4. Dual product line de-risked the business during the IPO setback
Cerebras’s hardware business (CS-3 sales to national labs and research institutions) provided a stable revenue base that allowed the company to continue investing in Cerebras Inference even when the IPO was blocked in late 2024. Without the hardware contracts — particularly the multi-year government and research institution engagements — the company would have faced more acute pressure to monetize the inference API faster, potentially forcing premature pricing moves. The dual-track structure gave management time to build inference API momentum while hardware revenue sustained operations. This separation of revenue streams with different predictability profiles proved to be a structural advantage.
Areas to improve : Gaps in Cerebras’s pricing and platform approach
1. No enterprise-grade access control or spend management on the API
Cerebras Inference lacks the organization-level cost controls enterprise buyers expect: team-based API key management, role-based access control, and spend caps or budget alerts per team or project. The console does now document usage, cached-usage and cost analytics plus admin-only audit logs — so the visibility layer exists — but nothing documented can actually stop spend. A company deploying Cerebras Inference across multiple engineering teams has no mechanism to prevent a single team from consuming the full month’s budget through an errant batch job. This cost unpredictability gap is a meaningful barrier to enterprise API adoption. Competitors like Anthropic and OpenAI offer organization dashboards with per-project spend limits and usage visibility. Until Cerebras adds these controls, it is effectively limited to developer-grade and small-team deployment patterns in the API tier.
2. Model catalog is too narrow for multi-model use cases
Cerebras’s public per-token rate card carries three models as of 2026-07-21 — GPT-OSS-120B, Google Deepmind Gemma 4 31B, and ZAI-GLM-4.7 — all open-source. The count understates how narrow the usable surface is: Gemma 4 31B arrived on 2026-07-21 classified Preview (evaluation only), and ZAI-GLM-4.7 is Preview and dated for deprecation on Aug 17, 2026, which leaves GPT-OSS-120B as the only model a buyer can put into production from the public card. Adding a model that the docs explicitly disclaim for production use widens the menu without widening what you can actually build on. A wider catalog (Llama, Qwen3, Mistral, DeepSeek, Kimi K2.x) exists only behind Dedicated Endpoints on custom pricing, so it is not self-serve. There is no access to Claude, GPT-4o, Gemini, or other frontier proprietary models through the Cerebras API. For organizations that want to consolidate their AI spend on a single inference platform, Cerebras cannot serve as a full-stack provider — it is a speed-optimized complement to, not a replacement for, platforms like Perplexity Sonar API or Fireworks AI. The narrow catalog also concentrates revenue risk on a small number of model relationships. Expanding the catalog — including through partnerships with model providers who would benefit from Cerebras’s speed for their open-weight models — is a strategic priority that has not yet materialized at scale.
3. Hardware pricing opacity creates a two-class developer ecosystem
The price opacity on CS-3 hardware creates an asymmetric developer experience: API users have a fully public rate card and can self-serve at any scale, while hardware prospects must engage sales before understanding costs. This bifurcation risks alienating the enterprise segment that Cerebras most needs to penetrate for long-term revenue growth. Publishing at least a floor price range or a “price per WSE-3 month” benchmark — even as a starting point for enterprise conversations — would help enterprises build preliminary business cases without requiring early sales engagement. Designing transparent enterprise pricing tiers is a solvable problem that Cerebras has not yet addressed.
Monetization stack & signals : how Cerebras builds & buys its revenue engine
Buys 1 Builds 0 2 open roles
Runs order-to-cash on NetSuite — a finance-led billing stack fit for lumpy hardware contracts, not the self-serve token API. Recent billing-close and FP&A hires point to formalizing revenue operations.
-
“You'll serve as a key link between Sales, Services, and Finance by leading month-end billing close activities, translating contract terms into billing in NetSuite, and managing credits and adjustments.”
- Director, Strategic Finance - Corporate FP&A RevOps seen May 21, 2026
- Billings and Collections Specialist Billing engineering seen Feb 13, 2026
Signals reviewed · derived from public job posts
Job postings fill and close over time — once a posting is filled we keep it as a dated citation (the quoted evidence remains); use View open roles for current listings.
Key takeaways
-
Speed can be a standalone value metric in a commoditized market. When every provider offers the same models at similar prices, a 10–20× throughput advantage is a genuine product differentiator that justifies category-level attention. Pricing teams competing in commoditized markets should ask what hardware or architectural advantages their infrastructure gives them that could be surfaced as a distinct pricing dimension.
-
Price-speed parity signals confidence; price premiums for speed signal weakness. Cerebras chose to price at parity with slower GPU-cloud competitors rather than charging a premium for faster throughput. This decision communicates confidence that speed adoption will be self-reinforcing — and it was. A speed premium would have created a price-performance comparison that competitors could close; price parity made the comparison entirely one-sided.
-
Free entry points should demonstrate the actual product — but they don’t have to be permanent. Cerebras has consistently exposed the same models at the same speed as the paid tier, so developers experience the core value proposition rather than a synthetic demo. What it changed on 2026-07-21 was the shape of the allowance: an open rate-limited tier became a $5 one-time credit. For usage-based pricing products, the durable rule is that you never degrade the quality of the free experience; the disposable part is whether the allowance renews. A credit grant is the tighter instrument — it gives finance a bounded, forecastable acquisition cost per signup, where an open tier’s cost scales with whoever shows up.
-
Dual hardware-and-API business models create resilience but complicate focus. Cerebras’s hardware revenue sustained the company through the IPO setback, but the hardware sales cycle (multi-year government contracts) and the API sales cycle (self-serve minutes) require fundamentally different GTM motions, pricing architectures, and customer success approaches. Companies with dual B2B models should be deliberate about keeping these tracks operationally separate.
-
Revenue concentration in a single international customer is an IPO-blocking risk. The G42 relationship — which represented a significant fraction of Cerebras’s hardware revenue — triggered the CFIUS review that blocked the IPO. For any AI infrastructure company serving international customers, the concentration risk from a single large non-US customer is not just a financial risk but a regulatory and liquidity risk that must be disclosed and managed from early stages.
UBP implications
-
Per-token pricing at high throughput creates a new cost efficiency frontier for real-time applications. Cerebras demonstrates that usage-based pricing in AI inference does not have to accept the latency/cost tradeoff as fixed. When a provider can deliver faster inference at the same price, it unlocks new use cases (real-time voice, streaming code generation, interactive document analysis) that were previously uneconomic on GPU clouds. UBP practitioners building AI products should evaluate whether speed-gated features are worth pricing separately — Cerebras’s experience suggests that speed, at parity pricing, drives adoption without requiring a premium.
-
Symmetric vs. asymmetric token pricing reflects maturity in cost modeling. Cerebras’s evolution from symmetric ($0.10/$0.10 for Llama 3.1 8B) to asymmetric pricing (Llama 3.3 70B at $0.85/$1.20, GPT-OSS-120B at $0.35/$0.75) maps directly to better understanding of actual prefill versus decode compute costs on their hardware. Choosing the right usage metric for inference APIs means understanding that output generation is fundamentally more compute-intensive than input processing — a nuance that takes real usage data to optimize in a rate card.
-
The free-to-paid gate moved from rate limits to credit exhaustion — and that changes who converts. Through mid-2026 Cerebras gated only on throughput: the full catalog was free, and rate limits were the value metric for the transition. Rate limits are a soft gate and a clean usage signal — anyone hitting them has a real workload, so the tier self-selects for deployment intent while curiosity costs the vendor a trickle. As of 2026-07-21 the gate is a $5 credit balance, which is a hard gate on a different axis: it stops on cumulative spend rather than on concurrency, so a slow-burning production trickle now converts (or churns) where it previously would have run free forever. Operators choosing between the two should be explicit about which behaviour they are pricing out — throttles push out intense users, credits push out persistent ones — and note that neither is free to reverse, since taking a standing free tier away is a visible takeaway in a way that tightening rate limits is not. See usage aggregation methods for how the underlying metering differs.
Sources
- Cerebras pricing page (tiers, per-token rate card, Cerebras Code plans) (accessed 2026-08-27)
- Cerebras Code product page (Free / Pro / Max plans, GLM 4.7) (accessed 2026-08-27)
- Cerebras Dedicated Endpoints — supported models (accessed 2026-08-27)
- Cerebras Inference model catalog (production vs preview) (accessed 2026-08-27)
- Cerebras docs — GPT OSS 120B model page (rates, context window, max output) (accessed 2026-07-21)
- Cerebras docs — Gemma 4 31B model page (rates, speed) (accessed 2026-07-21)
- Cerebras docs — Z.ai GLM 4.7 model page (rates, Aug 17 2026 deprecation) (accessed 2026-07-21)
- Cerebras docs — Rate limits (Free Trial / Developer / Enterprise tier terms) (accessed 2026-07-21)
- Cerebras docs — Account & Billing (free-credit terms and expiry) (accessed 2026-07-21)
- Cerebras docs — Usage & Monitoring (usage, cost, audit logs) (accessed 2026-07-21)
- Cerebras Inference pricing documentation (now redirects to cerebras.ai/pricing) (accessed 2026-07-21)
- Cerebras blog — GPT-OSS-120B runs fastest on Cerebras (accessed 2026-05-29)
- Cerebras Systems GitHub — Cloud SDK Python (accessed 2026-05-29)
- Cerebras Systems S-1 filing (August 2024) (accessed 2026-05-29)
Bottom line
Cerebras has built the fastest publicly available LLM inference platform on the market by removing Nvidia GPUs from the equation entirely — and has priced it at parity with far slower GPU-cloud competitors. The result is a genuinely novel value proposition: the same open-source models you can run anywhere else, at 10–20× the throughput, at the same or lower cost. The inference API is clean, OpenAI-compatible, and still cheap to try, though the July 2026 switch from an open free tier to a $5 credit trial means the evaluation window is now finite. The gaps are real but fixable — the public rate card lists three models and sanctions exactly one of them for production, enterprise controls are immature, and hardware pricing opacity limits self-serve enterprise evaluation. Cerebras is a compelling speed-optimized inference tier for organizations running high-throughput workloads on open-source models; it is not yet a full-stack AI platform.
Browse the full pricing blueprint to compare Cerebras against other AI infrastructure providers.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
Cerebras announces CS-4, a 4th-generation wafer-scale system
Cerebras announced CS-4, built from three new Wafer Scale Engine 3 Turbo processors on a redesigned modular rack ("Nexus Platform"). Cerebras claims up to 30x faster inference than GPU systems, up to 10x more throughput per watt and up to 2x faster performance than CS-3, and wafer-to-wafer interconnect latency as low as 2 microseconds. No pricing was disclosed — CS-4 remains an enterprise-contract, "contact us" product like CS-3 — with first shipments described as beginning "this quarter" (Q3 2026).
ZAI-GLM-4.7 deprecated and removed from the public rate card
Consistent with the deprecation date footnoted on the pricing page a month earlier, ZAI-GLM-4.7 ($2.25 input / $2.75 output per million tokens) was removed from Cerebras's public "Developer Tier Pricing" rate card and the docs Model Catalog. Confirmed via a fresh capture on 2026-08-26, which shows the rate card and Model Catalog both down to two models — GPT-OSS-120B (production) and Google Deepmind Gemma 4 31B (Preview). Z.AI's GLM 4.X and GLM 5.X model families remain reachable only through Dedicated Endpoints on custom reserved-capacity pricing; no per-token price moved on the two remaining public-card models.
Free tier becomes a $5 credit trial; Gemma 4 31B joins the rate card
Cerebras replaced its open, rate-limited Free tier with a Free Trial that grants $5 in one-time credits after account creation, ending the perpetual free path onto the platform. Docs terms tighten it further: the credits are issued only after adding a verified payment method, and they expire 30 days after grant. The public rate card grew back to three models with Google Deepmind Gemma 4 31B at $0.99 input / $1.49 output per million tokens (~1,800 tokens/s), classified Preview. ZAI GLM 4.7 kept its $2.25/$2.75 rates but picked up a footnoted deprecation date of Aug 17, 2026, leaving GPT OSS 120B ($0.35/$0.75) as the only production-sanctioned model on the public card. Cerebras Code Pro ($50/mo) and Max ($200/mo) were unchanged and still listed SOLD OUT.
Public rate card narrows; Cerebras Code subscriptions launch
By mid-2026 the public per-token rate card lists just two models — GPT-OSS-120B ($0.35/$0.75, production) and ZAI-GLM-4.7 ($2.25/$2.75, labeled a Preview/evaluation model). Llama and Qwen3 families moved to Dedicated Endpoints on reserved-capacity custom pricing. Access is now tiered (Free, a self-serve Developer tier from $10, and Enterprise), and Cerebras introduced fixed-price Cerebras Code coding plans — Pro at $50/month (24M tokens/day) and Max at $200/month (120M tokens/day), both sold out at launch.
Qwen-3-32B and ZAI-GLM-4.x Models Added
Cerebras expanded its model catalog with Alibaba's Qwen-3-32B (priced at $0.40/$0.80 per million tokens) and the ZAI-GLM-4.6 and 4.7 models from Zhipu AI (priced at $2.25/$2.75 per million tokens). ZAI-GLM-4.6 was subsequently deprecated in January 2026.
GPT-OSS-120B Added — Fastest Open Reasoning Model
Cerebras added OpenAI's open-source GPT-OSS-120B (Apache 2.0 license) to its inference cloud at $0.35 input/$0.75 output per million tokens, claiming the fastest inference speed for a 120B-class reasoning model. The model supports a 131K context window.
IPO Blocked by CFIUS National-Security Review
Cerebras's planned IPO (S-1 filed August 2024, targeting ~$8B valuation) was blocked when CFIUS opened a national-security review of the company's relationship with UAE-based G42, which held a significant revenue concentration and had prior ties to Huawei. The company withdrew the IPO registration.
Cerebras Inference Launched in Public Beta — 2,100 Tokens/Second
Cerebras launched Cerebras Inference as a public beta cloud API, delivering Llama 3.1 8B and 70B models at speeds of 2,100 and 450 tokens/second respectively — exceeding GPU-cloud alternatives by 20×. The launch included a free developer tier and usage-based pay-per-token pricing.
WSE-3 and CS-3 Announced — 4 Trillion Transistors
Cerebras announced the third-generation Wafer Scale Engine (WSE-3) with 4 trillion transistors and 900,000 cores, and the CS-3 compute system built around it. CS-3 is positioned for both training and inference at scale. Pricing remains enterprise-contract only.
Cerebras Model Studio — First Cloud API
Cerebras launched Cerebras Model Studio, an early cloud-based API giving customers access to GPT-J and other open-source models running on WSE hardware. This was the company's first foray into cloud inference, initially available only to existing hardware customers.
CS-2 System Launched with WSE-2
Cerebras launched the CS-2 compute system powered by the WSE-2 chip (2.6 trillion transistors, 850,000 cores, 40 GB SRAM). CS-2 was sold to national labs, healthcare systems, and enterprises for large-scale model training.
WSE-1 Unveiled at Hot Chips — First Wafer-Scale AI Chip
Cerebras unveiled the Wafer Scale Engine (WSE-1) at Hot Chips 2019: 1.2 trillion transistors, 400,000 AI-optimized cores, 18 GB on-chip SRAM on a 46,225 mm² die. The chip was sold as part of the CS-1 compute system for on-premises deep learning training.
Cerebras Systems Founded
Andrew Feldman and Gary Lauterbach founded Cerebras Systems in Los Altos, California, to build a purpose-built AI chip that would break through GPU memory bottlenecks by placing all SRAM on a single wafer-scale die.
- · Cerebras's Wafer Scale Engine 3 (WSE-3) contains 4 trillion transistors on a single silicon wafer — roughly 57× more transistors than Nvidia's H100 GPU — making it the largest chip ever manufactured as of 2024.
- · Cerebras filed for an IPO in August 2024 valuing the company at approximately $8 billion, but the IPO was blocked in November 2024 when the Committee on Foreign Investment in the United States (CFIUS) opened a national-security review related to the company's largest customer, G42 of the UAE, which had previously had ties to Huawei.
- · At launch in August 2024, Cerebras Inference ran Llama 3.1 70B at 2,100 tokens per second — more than 20× faster than GPU-based competitors like Together AI or Fireworks AI at the time, a speed record that attracted significant developer attention.
Questions & answers
- How much does Cerebras Inference cost per million tokens?
- The public Cerebras rate card lists two models as of August 2026: GPT-OSS-120B at $0.35 input/$0.75 output per million tokens (production) and Google Deepmind Gemma 4 31B at $0.99 input/$1.49 output (Preview). ZAI-GLM-4.7, previously listed at $2.25 input/$2.75 output, was removed from the public card on its disclosed deprecation date of August 17, 2026. Preview models are intended for evaluation only, so GPT-OSS-120B is the only public model sanctioned for production. Other models such as Llama 3.3 70B, Qwen3-32B, and now the Z.AI GLM family are available via Dedicated Endpoints on custom pricing rather than the public rate card.
- Does Cerebras offer a free tier for the inference API?
- No longer. As of July 21, 2026 Cerebras replaced its open free tier with a Free Trial that grants $5 in one-time credits after you create an account, with access to all Cerebras-powered models and Discord support. The docs add two conditions the pricing page leaves out: the credits are granted only after you add a verified payment method (adding one is free), and they expire 30 days after they are granted. That $5 is worth roughly 14 million input tokens on GPT-OSS-120B, and it does not renew. Past the trial, the self-serve Developer tier starts at just $10 and offers 10x higher rate limits and higher-priority processing; the Enterprise tier (contact sales) adds custom weights, dedicated queue priority, and guaranteed uptime.
- How fast is Cerebras inference compared to GPU-based providers?
- Cerebras Inference delivers 1,000–2,100 tokens per second on Llama 3.1 70B-class models, compared to 40–80 tokens/second on GPU-based providers like Together AI or Fireworks AI. The speed advantage comes from the on-chip SRAM of the WSE eliminating GPU memory bandwidth bottlenecks.
- What models are available on Cerebras Inference?
- As of August 2026, the public Cerebras per-token rate card lists just GPT-OSS-120B (production) and Google Deepmind Gemma 4 31B (Preview) — ZAI-GLM-4.7 was removed from the card on its disclosed August 17, 2026 deprecation date. A much wider catalog — including Llama 3.3 70B, Llama 4 Maverick and Scout, Qwen3-32B and Qwen3-235B, Qwen3-Coder, Mistral, DeepSeek, Kimi K2.x, and Z.AI's GLM 4.X/5.X families (including the former GLM-4.7) — is available through Dedicated Endpoints on reserved-capacity custom pricing. The catalog focuses on open-source models that benefit most from Cerebras's speed advantage.
- How does Cerebras hardware (CS-3, CS-4) pricing work?
- The CS-3 compute system is sold under enterprise contracts via a direct sales process. Pricing is not publicly listed and varies based on deployment configuration (on-premises, cloud-connected, or managed), support tiers, and commitment length. Typical CS-3 deployments involve multi-year contracts at research institutions and national labs. On August 18, 2026 Cerebras announced CS-4, a fourth-generation system built on three Wafer Scale Engine 3 Turbo processors, claimed up to 2x faster than CS-3 with first shipments beginning Q3 2026 — like CS-3, no public price is disclosed and access runs through the same "contact us" enterprise sales motion.
- Is Cerebras an alternative to Nvidia GPUs?
- For inference workloads on supported open-source models, yes. Cerebras Inference runs without Nvidia GPUs, delivering faster throughput at competitive per-token cost. For training arbitrary model architectures or running proprietary models, GPU-based infrastructure remains more flexible. Cerebras's hardware is optimized for dense transformer inference.