AI Summary
About
MiniMax (稀宇科技) is a Shanghai-based foundation-model company that builds proprietary multimodal models — text/reasoning (the MiniMax-M and legacy abab families), the open-weight MiniMax-H3 video model, Speech, and Music — and monetizes them across three distinct surfaces: consumer apps, monthly subscriptions, and a per-token developer API. The consumer side runs Talkie, an AI character role-play companion aimed at international markets, and MiniMax Hub, a multimodal creation app that showcases the video, speech, and music models and competes with OpenAI’s Sora (the app previously carried the Hailuo brand; MiniMax’s site no longer uses that name as of August 2026). The developer side exposes a pure usage-based API plus monthly Token Plan subscriptions for individuals and small teams.
Founded in early 2022 by ex-SenseTime researcher Yan Junjie (with Yang Bin and Zhou Yucong), MiniMax became one of China’s “AI tiger” labs, backed by Alibaba, Tencent, miHoYo, Hillhouse, HongShan, and IDG. In January 2026 it listed on the Hong Kong Stock Exchange, raising roughly US$619M at the top of its range (~$6.5B valuation) and surging about 43% on debut to a ~$9.3B valuation — making MiniMax a public company despite reporting only about $53M of revenue against a ~$512M loss for the first nine months of 2025. The raise funds the model and compute roadmap.
The strategic anchor of the price sheet is aggressive token economics paired with open weights. MiniMax-M1 (June 2025) shipped as the first open-source, large-scale, hybrid-attention reasoning model — weights on Hugging Face and GitHub, a 1M-token context window — and MiniMax-M2 (October 2025) is pitched explicitly as roughly 8% of Claude Sonnet 4.5’s token cost at nearly double the inference speed. Like Mistral AI, MiniMax couples open-weight releases with hosted per-token inference, but it leans harder on consumer media apps and ultra-low API pricing as its wedge.
Pricing summary : subscriptions, per-token API, and media units
MiniMax runs a three-surface model: freemium consumer apps, monthly subscriptions, and pure usage-based API pricing billed per million tokens. The dimensions are:
- Consumer apps — Talkie and MiniMax Hub are freemium with in-app credits/subscriptions; model use on the MiniMax app and web is free.
- Token Plan subscriptions — Plus ($22/mo), Max ($55/mo), Ultra ($132/mo), each also sold as a per-seat Token Plan for Teams license with a shared prepaid Credits pool. Tiers scale agent concurrency (3–7 simultaneous agents) and rolling/weekly quota windows rather than per-seat licensing, and all reach every API-platform model. Prices rose 10% across all three tiers as of the August 2026 snapshot (see Pricing evolution for the prior rates); a July 2026 revamp had doubled the included tokens at the prior prices, marketed at the time as “10× Claude Pro at the entry tier” and “up to ~12.5B tokens/month” — that specific marketing language is no longer shown on MiniMax’s live site, and no numeric token quota is currently published per tier.
- API tokens — separate input and output rates per million tokens by model (flagship MiniMax-M3 and MiniMax-M2.7 at $0.30 in / $1.20 out, cache reads $0.06; M3 above 512k context or M2.7-highspeed at $0.60 / $2.40; legacy M2/M2.1/M2.5 held at the same $0.30/$1.20 rate — open-weight M1 is no longer listed on the live pricing page).
- Media-gen units — MiniMax H3 video per second ($0.08 at 768P, $0.13 at 2K, replacing Hailuo’s per-clip pricing), Speech per million characters ($60–$100), Music 3.0 per track ($0.15), images per image ($0.0035).
- Tool & agent calls — MCP (API-vlm) and server tools (web_search) bill per call/request at $0.01 (API-vlm cut from $0.06 effective July 22, 2026), deducting from Token Plan quota then Credits.
- Prepaid credits — sold at 1,000 credits = $1, valid 365 days, as the common spend currency across the API.
What makes this different: MiniMax publishes raw per-million-token billing in USD on an international card at among the lowest frontier rates ($0.30/M input), and pairs it with open weights and per-clip media pricing — monetizing managed inference and consumer apps rather than the model itself.
Pricing by product
Token Plan — monthly subscriptions (USD)
| Tier | Price | Included | Key mechanics |
|---|---|---|---|
| Plus | $22 / mo | All API-platform models; 3–4 concurrent agents; 5-hour rolling + weekly quotas | Personal projects & prototyping |
| Max | $55 / mo | All models; 4–5 concurrent agents; higher quotas | Daily coding with agents + multimodal |
| Ultra | $132 / mo | All models; 6–7 concurrent agents; extended sessions | Heavy agent workflows |
Token Plan scales by agent concurrency and quota windows, not seats. Plus/Max/Ultra prices rose 10% as of the August 2026 snapshot (see Pricing evolution for the prior rates). Token Plan for Teams, now documented as its own surface on the international docs, lets a Team Owner buy the same Plus/Max/Ultra seats and assign them 1:1 to members — a subscription can be reassigned mid-cycle without resetting usage — plus a separate shared Credits pool that members without an assigned seat can also draw from if enabled; pay-as-you-go API keys stay on a per-team wallet, separate from Token Plan and Credits.
API — text & reasoning models (per million tokens, USD)
| Model | Input /M | Output /M | Key mechanics |
|---|---|---|---|
| MiniMax-M3 (≤512k ctx) | $0.30 | $1.20 | Flagship coding/agentic model (MSA, 1M context); “permanent 50% off” list $0.60 / $2.40; cache reads $0.06/M |
| MiniMax-M3 (>512k ctx) | $0.60 | $2.40 | Extended-context tier; cache reads $0.12/M |
| MiniMax-M2.7 | $0.30 | $1.20 | Agent/coding model; cache reads $0.06/M, writes $0.375/M |
| MiniMax-M2.7-highspeed | $0.60 | $2.40 | Faster-inference tier; cache reads $0.06/M, writes $0.375/M |
| MiniMax-M2 / M2.1 / M2.5 (legacy) | $0.30 | $1.20 | Held at legacy pricing after the M2.7/M3 launch |
| MiniMax-M2.5-highspeed / M2.1-highspeed (legacy) | $0.60 | $2.40 | Legacy high-speed tier |
M3 is now the flagship; M2, M2.1, and M2.5 (plus their highspeed variants) are held under “Legacy Models” at unchanged rates. Open-weight MiniMax-M1 is no longer listed on MiniMax’s live pay-as-you-go pricing page as of August 2026 (see Pricing evolution for its June 2025 launch pricing); its weights remain downloadable for self-hosting. Cache reads run $0.06/M on M3/M2.7; cache writes about $0.375/M. The China-native card quotes the same rates in RMB (M2 launch: ¥2.1 in / ¥8.4 out).
API — video generation (MiniMax H3, USD)
| Service | Price | Key mechanics |
|---|---|---|
| MiniMax-H3 video — 2K | $0.13 / second | Billed per second of generated output |
| MiniMax-H3 video — 768P | $0.08 / second | Billed per second of generated output |
| Input material — image | First 5 free, then $0.04 / additional image | Applies to image-conditioned generations |
| Input material — audio | Free | Applies to audio-conditioned generations |
| Input material — video | Billed by input duration at the output-resolution rate (2K $0.13/sec, 768P $0.08/sec) | Applies to video-conditioned generations |
| MiniMax-H3-Regeneration (768P → 2K) | $0.05 / second | Upscales a previously produced 768P video; billed per second of the regenerated output; first 5 input images free, then $0.025 each |
| MiniMax-H3-Context-IR | $0.90 / M in · $3.60 / M out | Separate per-token task-pricing tier of the H3 line |
MiniMax-H3, an open-weight, general-purpose, omni-modal generation model, replaced Hailuo as MiniMax’s video product and moved video billing from per-clip to per-second. Token Plan subscriptions do not currently cover MiniMax-H3 usage — it draws only from unrestricted Credits or pay-as-you-go billing even for subscribers.
API — audio, music & image (USD)
| Service | Price | Key mechanics |
|---|---|---|
| Speech 2.8 turbo | $60 / 1M characters | Text-to-speech |
| Speech 2.8 HD | $100 / 1M characters | Higher-fidelity TTS |
| Rapid voice cloning | $1.5 / voice | One-off clone |
| Voice design | $3.00 / voice | Synthetic voice creation |
| Music 3.0 (closed to new users) | $0.15 / up-to-5-minute track | RPM 120 (contact sales to raise); paid Music-3.0 API closed to new signups effective Aug 20, 2026 — existing paying users keep access. Music 2.6 held at the same $0.15 rate under the same closure |
| Lyrics generation (closed to new users) | $0.01 / song | Same Aug 20, 2026 new-user closure as Music 3.0/2.6; existing paying users keep access |
| image-01 | $0.0035 / image | Image generation |
Effective August 20, 2026, MiniMax closed the paid Music Generation and Lyrics Generation APIs to new users (existing paying users keep access) and discontinued the free music-generation APIs entirely (Music-3.0-free, Music-2.6-free, music-cover-free). MiniMax points affected users to the MiniMax Audio consumer app or the open-source MiniMax Music 3 model on Hugging Face instead.
API — tool & agent calls (USD)
| Service | Price | Key mechanics |
|---|---|---|
| MCP — API-vlm | $0.01 / call | Cut from $0.06/call effective July 22, 2026; called through Token Plan, it deducts from the included quota, then from purchased Credits |
| Server Tools — web_search | $0.01 / request | Beta; the model runs the search server-side and answers from the results |
Sales motions across products: PLG / self-serve for consumer apps, Token Plan subscriptions, and the entire pay-as-you-go API; enterprise volume is handled through the API platform’s prepaid voice/video packs at lower unit rates.
Hidden costs : What MiniMax users actually pay
MiniMax’s headline token rates are among the lowest in the frontier tier, but the real bill is shaped by three things the sticker doesn’t show: the 4× output-to-input ratio, the per-clip media units billed entirely outside the token meter, and the context-tier step-up that doubles rates past 512k (M3) or 200k input (M1). Two archetypes show how the total assembles.
Archetype 1 — a developer running a coding agent on the MiniMax-M2.7 API. A team running an autonomous coding agent at roughly 50M input + 15M output tokens/month, plus a Hailuo video pipeline generating 20,000 short clips for marketing.
| Line item | Monthly cost |
|---|---|
| MiniMax-M2.7 input — 50M tok @ $0.30/M | $15 |
| MiniMax-M2.7 output — 15M tok @ $1.20/M | $18 |
| Hailuo 2.3 video — 20,000 clips @ ~$0.30 | ~$6,000 |
| Estimated total | ~$6,033/mo |
The lesson: token inference is almost free at MiniMax’s rates — $33 for a month of heavy coding — but the media-gen units dominate the moment video enters the picture. Each Hailuo clip is priced per output, not per token, so a video-heavy workload swamps the LLM line by two orders of magnitude. Cache reads at $0.06/M further cut the (already tiny) token cost on repetitive agent context.
Archetype 2 — an indie developer on Token Plan Plus. One Plus subscription at $22/mo, occasionally topping up with prepaid credits for a burst of agent runs that exceed the rolling quota window.
| Line item | Monthly cost |
|---|---|
| Token Plan Plus | $22.00 |
| Prepaid credits top-up (est. 10,000 credits) | ~$10 |
| Estimated total | ~$32/mo |
Here the surprise is the rolling-and-weekly quota window: Plus caps concurrency at 3–4 agents and meters usage across 5-hour and weekly windows, so a burst of parallel agent runs can hit the ceiling mid-week and force a credits top-up (1,000 credits = $1) rather than a hard stop. The quota structure, not a seat count, is the real cost lever.
Want to estimate your own MiniMax bill? Use the MiniMax pricing calculator to model your costs based on token volume, video clips, and subscription tier.
Pricing evolution : MiniMax pricing history and changes
MiniMax’s pricing evolved from app-only consumer monetization toward a published, ultra-low per-token API and a tiered subscription ladder, capped by a public-market listing. The API side has anchored on aggressive token economics since the M1 open-weight release; the consumer side ran on credits and app subscriptions from the start. The dated milestones below are reconstructed from primary launch posts and contemporaneous press.
Cadence
| Quarter | Price changes | Product / SKU additions | Notes |
|---|---|---|---|
| 2022 Q4 | 0 | 1 | MiniMax founded; abab LLM family + Talkie companion app |
| 2024 Q1 | 0 | 1 | Hailuo AI consumer multimodal platform launches |
| 2024 Q3 | 0 | 1 | Hailuo video-01 model ships (Sora competitor) |
| 2025 Q2 | 1 | 1 | 2025-06-16 MiniMax-M1 open-weight reasoning + length-tiered API pricing |
| 2025 Q4 | 1 | 1 | 2025-10-27 MiniMax-M2 at $0.30/M in, $1.20/M out; Coding/Agent plans |
| 2026 Q1 | 0 | 1 | 2026-01 Hong Kong IPO raises ~$619M; ~$9.3B debut valuation |
| 2026 Q2 | 0 | 1 | 2026-05-31 MiniMax-M3 ships as new flagship (frontier coding, MSA 1M context, multimodal) |
| 2026 Q3 | 2 | 3 | 2026-07 M2 line → M2.7 (+highspeed); Music 3.0 launches; MCP + server-tool per-call billing added; API-vlm cut $0.06→$0.01/call. 2026-08 Token Plan +10% across Plus/Max/Ultra; Token Plan for Teams launches; Music/Lyrics Generation APIs close to new signups (Aug 20) |
Tracked range: 2022 Q4–2026 Q3. Quarters not listed had no publicly announced price or SKU change. Dated milestones below cite primary launch posts and press.
Notable changes
- 2022 (early) — MiniMax founded in Shanghai; abab LLM family and the Talkie companion app monetize via app subscriptions and in-app credits.
- 2024-03 — Hailuo AI consumer platform launches, showcasing video/speech/music models on a freemium-plus-credits model.
- 2024-09 — Hailuo video-01 ships as a Sora competitor; video becomes a per-clip billable surface.
- 2025-06-16 — MiniMax-M1 launches open-weight (1M context) with length-tiered API pricing: $0.40/M in, $2.20/M out up to 200k, then $1.30/M in to 1M — the first tier undercutting DeepSeek-R1 (MiniMax M1 post).
- 2025-10-27 — MiniMax-M2 launches at $0.30/M input ($2.1 RMB) and $1.20/M output ($8.4 RMB), pitched at roughly 8% of Claude Sonnet 4.5’s cost, with a Coding Plan and Agent Plan (MiniMax M2 post).
- 2026-01 — MiniMax lists in Hong Kong, raising ~$619M and surging ~43% on debut to a ~$9.3B valuation (reported by Reuters, WinBuzzer, Yahoo Finance).
- 2026-05-31 — MiniMax-M3 ships as the new flagship: frontier coding/agentic, MSA sparse attention, 1M context, natively multimodal, priced at $0.30/M input (a “permanent 50% off” the $0.60 list) (MiniMax news).
- 2026-07 — Model lineup refreshed on the international card: the branded M2 line becomes MiniMax-M2.7 (plus an M2.7-highspeed tier at $0.60/$2.40), M1/M2/M2.5 move to “Legacy Models,” Music 3.0 launches at $0.15/track, and MCP (API-vlm) + server-tool (web_search) per-call billing appears — with API-vlm cut from $0.06 to $0.01/call effective July 22, 2026. Token Plan prices held at $20/$50/$120 while a revamp roughly doubled the included tokens.
- 2026-08-20 — MiniMax closes its paid Music Generation and Lyrics Generation APIs to new signups (existing paying users keep access at the same $0.15/track and $0.01/song rates) and retires the free music-generation tiers (Music-3.0-free, Music-2.6-free, music-cover-free) outright, pointing affected users to the MiniMax Audio consumer app or the open-source MiniMax Music 3 weights on Hugging Face.
- 2026-08 — MiniMax raises Token Plan subscription prices 10% across the board — Plus $20→$22, Max $50→$55, Ultra $120→$132 — with agent-concurrency limits and quota windows unchanged, and ships Token Plan for Teams, a per-seat licensing surface (1:1 seat assignment, mid-cycle reassignment without a usage reset, plus a shared Credits pool) layered on the same Plus/Max/Ultra tiers.
The open-weight + ultra-low-price wedge in detail
MiniMax’s pricing arc is a single bet: drive token cost toward the floor while open-sourcing the reasoning model that does the work. M1 was given away as weights and priced at $0.40/M input; M2 pushed the hosted rate to $0.30/M and benchmarked it explicitly against Claude as “8% of the cost.” The July 2026 refresh kept the same posture: the branded line rolled M2→M2.7 and the Token Plan revamp roughly doubled included tokens at unchanged $20/$50/$120 prices (re-anchored as “$20 = 10× Claude Pro”), while the newly opened MCP tool-call surface launched already discounted — API-vlm cut 6× to $0.01/call from July 22. Every surface MiniMax adds is priced at the floor rather than to defend margin. The pricing implication is that MiniMax does not expect to monetize the model directly — app subscriptions, media-gen units, and managed inference at razor-thin margins are the revenue, while the open weights and headline rates are distribution. The January 2026 IPO on roughly $53M of revenue, followed by a further ~$2B share-and-bond raise in July 2026 (reported by SiliconANGLE and Crypto Briefing), capitalizes that land-grab rather than a profitable book — a structurally different posture from Western labs charging $15–$75 per million output tokens on flagship models.
The floor isn’t held everywhere at once, though. The August 2026 Token Plan repricing — Plus/Max/Ultra all up 10%, to $22/$55/$132 — left the per-token API rate (the number MiniMax uses to sell against Claude) untouched and raised the packaged-subscription price instead. Floor pricing stays concentrated on the surface that does the marketing work; the surface that converts into recurring revenue is where margin gets defended first.
What’s unique : MiniMax’s distinctive pricing mechanics
1. Floor-seeking token economics, benchmarked against the West — but not on every surface. MiniMax prices M2.7/M3 at $0.30/M input — and explicitly markets it as “8% of Claude Sonnet 4.5’s cost.” Rather than competing on capability narrative, it competes on a published unit price that anchors against a named Western flagship, turning the per-token rate itself into the differentiator. In July 2026 MiniMax cut the new MCP (API-vlm) tool-call rate 6× from $0.06 to $0.01/call and doubled the tokens included in each Token Plan tier at the same price. But by August 2026 that instinct stopped being universal: MiniMax raised Plus/Max/Ultra Token Plan prices 10% (to $22/$55/$132) while leaving the per-token API rate untouched — the packaged-subscription surface is now where MiniMax defends margin, even as the raw API rate stays the floor-priced anchor against Claude.
2. Subscriptions metered by agent concurrency, not seats — with a seat wrapper for teams. Token Plan tiers (Plus/Max/Ultra) scale by how many agents you can run simultaneously (3–7) and by rolling/weekly quota windows — not by per-user licensing — for an individual buyer. The August 2026 Token Plan for Teams surface layers per-seat licensing (1:1 assignment, mid-cycle reassignment without a usage reset, plus a shared Credits pool) on top of the same concurrency tiers for team accounts, but the underlying meter doesn’t change: teams buy N concurrency-metered seats rather than a new per-user metric. For an agentic workload, the value metric is still parallelism and throughput, a closer proxy to value than a seat count alone.
3. Three monetization surfaces over one model stack. The same underlying models power consumer apps (Talkie, Hailuo AI) with credits, monthly Token Plan subscriptions, and a per-token API — three packagings of one stack. Media generation is priced per output unit (per clip, per character, per track) outside the token meter, so the meter shifts toward outcome-shaped units the moment a workload becomes multimodal. The July 2026 addition of per-call tool billing (MCP API-vlm and server-tool web_search at $0.01/call) extends that pattern to agent actions — priced per call, drawn down from the same Token Plan quota, so tool use and token use share one meter.
Strengths & weaknesses
| Strengths | Weaknesses |
|---|---|
| Among the lowest published frontier token rates ($0.30/M input on M2/M3) | Output rate is 4× input on M2/M3, and legacy tiers (M2/M2.1/M2.5) hold the same rate — but the open-weight M1 tier is no longer listed at all, and output-heavy jobs still cost more than the $0.30 headline implies |
| Open weights (M1) let buyers self-host the reasoning model — a credible lock-in hedge | Per-clip Hailuo media units ($0.19–$0.56) can dwarf the token bill on multimodal workloads, harder to predict than the LLM rate |
| One model stack monetized three ways (apps, subs, API) widens the funnel | China/international split (RMB vs USD cards) and frequent model renames (M1→M2→M2.7→M3) complicate price tracking |
| Subscriptions metered by agent concurrency map better to agentic value than seats — and Token Plan for Teams (Aug 2026) keeps that meter even when the buyer is a team | Token Plan quota windows (5-hour rolling + weekly) are concurrency caps, not published token quotas — opacity on the subscription side, and Teams’ per-seat purchase unit still reads as “seats” to a buyer comparing vendors |
| Cache reads at $0.06/M cut repetitive-context cost on agent loops | Public-company economics: ~$53M revenue against ~$512M loss signals the low prices are a land-grab, not yet a sustainable margin |
| USD international card publishes raw rates — no “contact sales” wall for inference | M3 above 512k context doubles to $0.60/$2.40, and the high-context tier is “limited availability” pending public release |
| A uniform 10% Token Plan repricing (Aug 2026) hit all three tiers alike, not just the entry tier — a cleaner signal of subscription-side demand than a single-tier squeeze | That same August 2026 increase cracks the pure floor-seeking story: subscriptions now defend margin even as per-token API rates hold at the floor, and it landed with no dedicated changelog note for renewing subscribers |
Billing UX : usage tracking and overage controls
- Prepaid credits as common currency — the API platform sells credits at 1,000 = $1 (valid 365 days), used as the spend unit across token, media, and audio consumption.
- Rolling + weekly quota windows — Token Plan subscriptions meter usage across a 5-hour rolling window and a weekly window rather than a flat monthly token bucket, smoothing burst usage.
- Cache-hit pricing — repeated context is billed at the cheaper cache-read rate ($0.06/M on M3/M2.7), letting agent loops with stable system prompts cut cost automatically.
- Tool calls draw down Token Plan quota — MCP (API-vlm) and server-tool (web_search) calls, billed $0.01 each pay-as-you-go, deduct from the included Token Plan quota first and spill to purchased Credits, so agent tool use shares one meter with tokens.
- Token Plan quota excludes MiniMax-H3 — the monthly Plus/Max/Ultra quota covers the full M3/M2.7/image/speech/music lineup, but MiniMax-H3 video, voice design, and rapid voice cloning are explicitly “not currently supported” under the plan quota — that usage draws only from unrestricted Credits or pay-as-you-go billing, even for subscribers.
- Prepaid voice/video packs — Audio Subscription sells prepaid audio-point packages (5 tiers, monthly with 3-/12-month prepay discounts) at lower unit rates than pay-as-you-go, covering all speech models. A separate “Video Packages” surface sells prepaid video-point packages, but as of August 2026 it explicitly supports only the legacy Hailuo model line — MiniMax-H3 usage is not yet eligible for prepaid packages and still draws from Credits or pay-as-you-go billing.
- Dual currency cards — an international USD card (platform.minimax.io) and a China-native RMB card (platform.minimaxi.com) expose the same products with separate key systems.
- Free app/web tier — model use inside the MiniMax app and web is free, serving as the top-of-funnel before API/subscription monetization.
- Token Plan for Teams — a Team Owner buys Token Plan seats (same Plus/Max/Ultra prices) assigned 1:1 to members; reassigning a seat mid-cycle keeps the current usage state rather than resetting it. A separate shared Credits pool lets members without an assigned seat still draw usage if the Owner/Admin enables Credits access for them, and each Team keeps its own pay-as-you-go wallet, kept separate from Token Plan and Credits.
- Grandfathered API closures — MiniMax closes paid APIs to new signups while keeping existing paying customers on the old terms, as with the Music/Lyrics Generation APIs closed to new users on August 20, 2026 (free music-generation tiers were retired outright, with no grandfathering).
Strategic wins : Why MiniMax’s pricing decisions worked
1. Pricing against a named competitor, not in a vacuum
By marketing M2 as “8% of Claude Sonnet 4.5’s cost,” MiniMax gave buyers a single, memorable price story anchored to a flagship they already know. It reframes the purchase from “is this model good enough?” to “why pay 12× more?” — a clean wedge for cost-sensitive developers. This is the inverse of capability-led pricing and mirrors the shift toward value-anchored, comparative pricing.
2. Open weights as distribution, inference as revenue
Open-sourcing M1 (1M context, on Hugging Face/GitHub) turned the reasoning model into a marketing asset while monetizing hosted inference and consumer apps. Developers evangelize the open weights; MiniMax earns on managed inference, media units, and subscriptions. Metering the delivery rather than the artifact is the same durable move covered in usage-based pricing strategy.
3. Concurrency-based subscriptions for an agentic world
Pricing Token Plan by simultaneous-agent capacity (3–7) and quota windows — not seats — aligns the meter with how agentic workloads actually consume compute. As one user can drive many parallel agents, seats stop tracking value; concurrency does. Choosing a usage metric that tracks parallelism is a forward-looking bet most subscription vendors haven’t made. MiniMax’s August 2026 Token Plan for Teams answers the obvious follow-up — how do teams buy concurrency? — by selling per-seat licenses of the same concurrency tiers rather than inventing a new metric, keeping the underlying meter consistent even as the purchase unit becomes a seat for multi-person accounts.
Areas to improve : Gaps in MiniMax’s pricing approach
1. Publish concrete token quotas, not just concurrency caps
Token Plan tiers state agent concurrency and “rolling/weekly quota windows” but not numeric token allowances. That opacity invites the bill-shock and unpredictability anxiety subscriptions are meant to remove. A concrete per-tier token quota — even approximate — would let buyers self-select without fear of silent throttling mid-week.
2. Surface media-gen cost alongside the token rate
Hailuo video, Speech, and Music are billed per output unit entirely outside the token meter, and for multimodal workloads they dominate the bill. Headlining only the $0.30/M token rate understates true cost. A combined “estimated cost per generation run” view would make multimodal totals predictable before commitment.
3. Reconcile the RMB and USD cards
The China-native RMB card and the international USD card expose the same products with separate key systems and occasionally diverging tiers. A single canonical price table with a currency toggle (as Mistral AI does) would reduce confusion for global buyers comparing the two surfaces.
4. Announce subscription price increases explicitly
The August 2026 Token Plan repricing (Plus/Max/Ultra all +10%) showed up only in a routine pricing-page capture, without a dedicated announcement post pairing it with a change list the way the M2.7/M3 launches and the July 2026 quota revamp got. Renewing subscribers have no changelog entry to point to for why their bill moved. A short “what changed and why” note — even a one-line changelog entry — would match the transparency MiniMax already extends to its per-token API rate card.
Monetization stack & signals : how MiniMax builds & buys its revenue engine
Buys 3 Builds 0
Bought its entire international revenue engine off the shelf — Stripe Payments, Invoicing and Tax across 100+ countries — while the China-native card stays on separate domestic rails. The in-house build is the model and token meter behind the price, not the billing plumbing.
-
“MiniMax implemented Stripe Payments to process first-time and recurring payments for its subscribers. With Payments, MiniMax immediately began collecting revenue from users in more than 100 countries using a single payments infrastructure.”
-
“MiniMax began using Stripe Invoicing to automate invoice generation and management. The team replaced error-prone Word documents with Stripe's automated invoicing tools that enable localized formats.”
-
“To manage its tax compliance requirements, MiniMax also added Stripe Tax to automate end-to-end tax management.”
Signals reviewed · derived from press & filings
Key takeaways
- Anchor the price to a named rival. “8% of Claude Sonnet 4.5” is a sharper sales tool than any benchmark chart — comparative unit pricing reframes the buying decision around cost, not capability.
- Give away the model, sell the delivery. Open-sourcing M1 turned R&D into distribution; revenue lives in hosted inference, consumer apps, and media units. Meter the delivery, not the artifact.
- Meter agents by concurrency, not seats. When one user runs many parallel agents, a seat count stops tracking value. Subscriptions priced by simultaneous-agent capacity map closer to consumption.
- Multimodal shifts the meter to outcome units. Per-clip video and per-character speech dominate the bill the moment a workload goes beyond text — the variable cost moves from tokens toward output units.
- Low prices can be a land-grab, not a margin — and the floor isn’t held everywhere at once. A January 2026 IPO on ~$53M revenue against a ~$512M loss shows MiniMax is buying share with floor-seeking prices and capitalizing the gap on public markets; the August 2026 10% Token Plan increase, paired with an untouched per-token API rate, shows that floor gets defended selectively once a surface starts converting into recurring revenue.
UBP implications
- Comparative pricing is a usage-pricing tactic — and often the surface that stays flat. Pricing a meter explicitly as a fraction of a named competitor’s rate (“8% of Claude”) turns the unit price into the headline value metric; MiniMax held that per-token rate steady even while raising Token Plan subscription prices 10% in August 2026, because the anchor rate does the marketing work the packaged tier doesn’t. UBP strategists should consider anchoring rates to a reference competitor, not just to cost — and expect that anchor surface to move independently of packaged/subscription pricing.
- The meter migrates from tokens to output units as models go multimodal. MiniMax bills text per token but video per clip and speech per character — an early signal that outcome-shaped pricing emerges naturally once a single stack spans modalities. UBP design should anticipate per-output units alongside tokens.
- Concurrency can be a cleaner value metric than seats for agents. As one user drives many parallel agents, parallelism tracks value better than user count. Practitioners building usage-based subscriptions for agentic products should evaluate concurrency and quota windows as the priced dimension.
Sources
- MiniMax pricing overview (accessed 2026-08-04) —
www.minimax.io/pricenow redirects to the MiniMax homepage;www.minimax.io/pricingredirects here - MiniMax API pay-as-you-go pricing (accessed 2026-08-04)
- MiniMax Token Plan subscriptions (accessed 2026-08-04)
- MiniMax API release notes (accessed 2026-08-04)
- MiniMax Video Packages pricing (accessed 2026-08-04)
- MiniMax Audio Subscription pricing (accessed 2026-08-04)
- MiniMax-M1 launch announcement (accessed 2026-08-04)
- MiniMax-M2 launch announcement (accessed 2026-06-11)
- Artificial Analysis — MiniMax-M2 model page (accessed 2026-08-26) — independent third-party confirmation of the $0.30/M in, $1.20/M out base API rate
- OpenRouter — MiniMax models (accessed 2026-08-26) — independent third-party confirmation of M3/M2.1 API rates, H3 video ($0.13/sec, 2K), and Speech 2.8 turbo/HD rates
- Browse the pricing blueprint corpus
Bottom line
MiniMax prices three surfaces from one model stack: freemium consumer apps (Talkie, MiniMax Hub), monthly Token Plan subscriptions ($22–$132/mo metered by agent concurrency, raised 10% in August 2026 and now also sold per-seat through Token Plan for Teams), and a pure per-token API among the lowest in the frontier tier ($0.30/M input on the current M3/M2.7 line, with M2/M2.1/M2.5 held as legacy and open-weight M1 no longer listed on the live pricing page), plus per-second MiniMax-H3 video, per-character speech (with the paid Music/Lyrics APIs closed to new signups as of August 20, 2026), and — since July 2026 — per-call MCP and web_search tool billing at $0.01 each. The open-weight M1 and explicit “8% of Claude” positioning still make the raw API rate the differentiator; the August subscription increase shows the floor-seeking instinct hasn’t extended to the packaged surface. The January 2026 Hong Kong IPO and a further ~$2B July raise capitalize a land-grab on ~$53M of revenue. The main friction is concurrency-only quota opacity and media units that swamp the token bill on multimodal workloads.
Want to compare MiniMax against other foundation-model providers? See Mistral AI, or browse the full pricing blueprint.
Pricing timeline : Major events on a vertical axis
Each milestone below corresponds to a public pricing change, product launch, or material adjustment. Major events use a filled marker; minor adjustments use a faded one.
Token Plan raised 10% across Plus/Max/Ultra
MiniMax raises Token Plan subscription prices 10% on the international USD card — Plus $20→$22, Max $50→$55, Ultra $120→$132 per month — with agent-concurrency limits, quota windows, and per-token API rates unchanged. The increase applies uniformly across all three tiers rather than to a single tier, first observed in the August 2026 capture.
Token Plan for Teams launches; Music/Lyrics APIs close to new signups
MiniMax documents Token Plan for Teams as its own pricing surface — per-seat Plus/Max/Ultra licenses with 1:1 assignment, mid-cycle reassignment, and a shared Credits pool for unassigned members — while closing the paid Music and Lyrics Generation APIs to new users effective August 20, 2026 (existing paying users keep access at the same rates) and retiring the free music-generation tiers outright.
Lineup refresh: M2→M2.7, Music 3.0, MCP + server-tool billing; API-vlm cut to $0.01/call
Live USD international card: flagship MiniMax-M3 and MiniMax-M2.7 at $0.30/M in / $1.20/M out (cache read $0.06), M3 >512k and M2.7-highspeed at $0.60 / $2.40; M1/M2/M2.5 moved to 'Legacy Models.' Music 3.0 launches at $0.15/track (free tier at RPM 3). New per-call surfaces: MCP API-vlm cut $0.06→$0.01/call (effective 2026-07-22) and server-tool web_search at $0.01/request, both deducting from Token Plan quota then Credits. Token Plan held at $20/$50/$120 with a '2× tokens' revamp (up to ~12.5B tokens/month).
Live snapshot: Token Plan $20–$120, M2/M3 API $0.30/M, Hailuo media units
Captured live USD international card: Token Plan subscriptions Plus $20 / Max $50 / Ultra $120 per month; per-token API MiniMax-M2/M3 at $0.30 in / $1.20 out (cache read $0.06), M3 >512k at $0.60 / $2.40; media-gen units — Hailuo 2.3 video $0.19–$0.56/clip, Speech 2.8 $60–$100/M chars, Music 2.6 $0.15/track, image-01 $0.0035/image; credits at 1,000 = $1.
MiniMax-M3 ships as new flagship (frontier coding, 1M context, multimodal)
MiniMax releases M3, a frontier coding/agentic model built on its novel MSA (MiniMax Sparse Attention) supporting up to 1M context and natively multimodal — positioned as the only domestic model combining all three 'Frontier essentials' and the only open-source one in its class. Priced at $0.30/M input on the API (a 'permanent 50% off' the $0.60 list), it becomes the headline model above the M2 line. (Source: MiniMax M3 release post, 2026-05-31.)
Hong Kong IPO raises about $619M
MiniMax lists on the Hong Kong Stock Exchange in January 2026, raising roughly HK$4.8B (about US$619M) priced at the top of its range (~$6.5B valuation), then surging about 42.7% on debut to a ~$9.3B valuation. Cornerstone investors include Abu Dhabi's ADIA; backers include Alibaba, Tencent, miHoYo, Hillhouse, and IDG. The raise funds the model and compute roadmap. (Source: HK IPO press, 2026-01.)
MiniMax-M2 launches at $0.30/M in, $1.20/M out
MiniMax releases M2, an agent/coding-focused model priced at $0.30/M input ($2.1 RMB) and $1.20/M output ($8.4 RMB) — pitched as roughly 8% of Claude Sonnet 4.5's token cost at nearly double the inference speed. A 197k context window, cache-hit pricing, and a Coding Plan / Agent Plan accompany the launch. (Source: MiniMax M2 launch post, 2025-10.)
MiniMax-M1 open-weight reasoning model + length-tiered API pricing
MiniMax ships MiniMax-M1 — billed as the first open-source, large-scale, hybrid-attention reasoning model — with weights on Hugging Face and GitHub, a 1M-token context window, and 80k-token reasoning output. API pricing is tiered by input length: $0.40/M in and $2.20/M out up to 200k tokens, then $1.30/M in and $2.20/M out to 1M. The first tier undercuts DeepSeek-R1; app and web use are free. (Source: MiniMax M1 launch post, 2025-06.)
Hailuo video-01 model ships (Sora competitor)
MiniMax releases video-01, its text-to-video model, positioning Hailuo against OpenAI's Sora. Video generation becomes a distinct billable surface, later priced per clip on the API by resolution and duration.
Hailuo AI consumer platform launches
MiniMax launches Hailuo AI, a consumer multimodal platform that becomes the showcase for its video, speech, and music models. Consumer access is freemium with credit/subscription monetization inside the app, separate from the developer API.
MiniMax founded; abab LLM family and consumer apps
MiniMax is founded in Shanghai in early 2022 by ex-SenseTime researcher Yan Junjie, building its proprietary abab large-language-model family and launching consumer apps including the Talkie AI character companion for international markets — monetized via app subscriptions and in-app credits rather than a public API price sheet at first.
- · MiniMax-M1 (June 2025) was billed as the first open-source, large-scale, hybrid-attention reasoning model — open weights on Hugging Face, a 1M-token context, and an 80k-token reasoning budget.
- · MiniMax pitches M2's $0.30/M input price as roughly 8% of Claude Sonnet 4.5's token cost at nearly twice the inference speed.
- · MiniMax raised about $619M in its January 2026 Hong Kong IPO and jumped about 43% on debut to a ~$9.3B valuation — on just ~$53M of revenue against a ~$512M loss for the first nine months of 2025.
Questions & answers
- What is MiniMax's pricing model?
- MiniMax runs a three-surface model: free/credit-based consumer apps (Talkie, MiniMax Hub), monthly Token Plan subscriptions (Plus $22, Max $55, Ultra $132), and a pure per-token API billed per million tokens (MiniMax-M2/M3 from $0.30 in / $1.20 out).
- How much does the MiniMax API cost per million tokens?
- The flagship MiniMax-M3 and the MiniMax-M2.7 model cost $0.30 per million input tokens and $1.20 per million output tokens, with cache reads at $0.06/M. M3 above 512k context (and the M2.7-highspeed tier) is $0.60 in / $2.40 out. Legacy models M2/M2.1/M2.5 hold the same $0.30/$1.20 rate; the open-weight M1 launched at length-tiered pricing ($0.40–$1.30 in / $2.20 out) in June 2025 but is no longer listed on MiniMax's live pay-as-you-go pricing page.
- How much are MiniMax Token Plan subscriptions?
- Token Plan is $22/mo (Plus), $55/mo (Max), and $132/mo (Ultra) — raised 10% from $20/$50/$120 as of August 2026. Tiers scale agent concurrency (3–7 agents) and rolling/weekly quota windows rather than per-seat licensing, and all tiers reach every model on the API platform. A separate Token Plan for Teams surface sells the same tiers as per-seat licenses with a shared Credits pool.
- How is MiniMax video and speech priced?
- MiniMax-H3, the open-weight video model that replaced Hailuo in mid-2026, bills per second of generated output: $0.08/second at 768P or $0.13/second at 2K, plus a $0.05/second regeneration rate and per-token H3-Context-IR pricing ($0.90/M in, $3.60/M out). Speech 2.8 is $60/M characters (turbo) or $100/M (HD), and Music 3.0 is $0.15 per up-to-5-minute track — but effective August 20, 2026, MiniMax closed the paid Music and Lyrics Generation APIs to new signups (existing paying users keep access) and retired the free music-generation tiers outright.
- Is MiniMax-M1 open weight?
- Yes. MiniMax-M1, released June 2025, was the first open-source large-scale hybrid-attention reasoning model, with weights on Hugging Face and GitHub, a 1M-token context window, and unlimited free use in the MiniMax app and web. Its hosted per-token API pricing is no longer listed on MiniMax's live pricing page as of August 2026, but the open weights remain downloadable for self-hosting.
- Is MiniMax pricing in USD or RMB?
- The international card (platform.minimax.io / minimax.io) is in USD; the China-native card (platform.minimaxi.com) is in RMB. M2 launched at $0.30/M input (¥2.1) and $1.20/M output (¥8.4).