Zero-markup resale transmits a model cut in days
Between 2026-08-04 and 2026-08-11, five corpus companies — Glean, Cursor, Augment Code, LiveKit and Vercel — repriced the same two OpenAI models by exactly the same two percentages: GPT-5.6 Luna down 80%, GPT-5.6 Terra down 20%. None of them made a pricing decision: each already billed third-party model tokens at cost, so the upstream move passed straight through. The second measured transmission was slower — OpenAI's GPT-5.6 Sol cut took 12 to about 20 days to reach most resellers — so 'in days' now means days to weeks.
What's happening — and why
What's happening: OpenAI launched the GPT-5.6 line on 2026-07-23 — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens. Glean's Model Hub table then cut Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) and Terra 20% ($2.50 to $2.00, $15.00 to $12.00) on 08-04. Cursor matched on 08-06, and Augment Code, LiveKit and Vercel all matched on 08-11. Netlify ran the identical mechanic on a different upstream three days later, halving Gemini 3.6 Flash ($1.50/$7.50 to $0.75/$3.75 per 1M) at a fixed 180-credits-per-dollar conversion, with no plan change at all (Free $0 / Personal $9 / Pro $20 unchanged).
Why: each of these vendors had already published an at-cost passthrough policy before the window opened. Cursor's docs state that its Other Models pool bills third-party API pricing at cost; Vercel attributes its move to the zero-markup AI Gateway; Netlify converts provider rates at a fixed 180 credits per $1. Having ceded pricing authority for third-party tokens, they don't decide these numbers — they inherit them. Their model prices are consequences, not decisions.
What makes it hard to see: the denomination changes at each hop. Glean, Cursor and Augment quote per 1M tokens; LiveKit quotes per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%; Netlify quotes in credits. Vercel's v0 Mini fell from $1/$5 to $0.20/$1.20 — landing on Luna's exact new rate — yet the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all. And the direction of the underlying move matters: Luna launched at $1/$6, above the outgoing GPT-5.4 mini at $0.75/$4.50, so the 80% cut twelve days later was a correction to a launch price. Every reseller inherited both moves.
What September showed: the lag varies. OpenAI's 2026-08-26 GPT-5.6 Sol cut to $4/$20 reached Cursor the same day, but Netlify after 12 days (2026-09-07), Augment Code after 15 (2026-09-10) and Glean's card after about 20; AssemblyAI's LLM Gateway carried the Luna and Terra cuts 25 days after OpenAI's own change (2026-09-20). The transmitted price can also carry an expiry the reseller does not print: OpenAI labels the Sol rate promotional through 2026-11-21. And the markup is being renamed rather than removed at the gateway layer — LangChain's LLM Gateway sells tokens 'at face value' with a 3.5% processing fee on credit purchases, and OpenRouter added an 8% Business tier beside its 5.5% Standard fee.
How it works
Evidence over time
19 supporting · 0 counter — hover or tap a point for detail, click to jump to the row.
Evidence
| Company | Date | What happened |
|---|---|---|
| OpenAI | Jul 2026 | Launched the GPT-5.6 line — Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per 1M tokens — replacing the GPT-5.5/5.4 headline lineup. Luna at $1/$6 sat ABOVE the outgoing GPT-5.4 mini at $0.75/$4.50, i.e. the everyday tier launched more expensive. |
| Glean | Aug 2026 | Model Hub Usage table cut GPT-5.6 Terra 20% ($2.50 to $2.00 in, $15.00 to $12.00 out) and GPT-5.6 Luna 80% ($1.00 to $0.20 in, $6.00 to $1.20 out) — the first published price decreases on that card since Glean began disclosing dollar rates in July 2026. Luna was simultaneously reclassified Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits. |
| Cursor | Aug 2026 | Other Models pool: Luna $1/$1.25/$0.10/$6 to $0.20/$0.25/$0.02/$1.20; Terra $2.50/$3.125/$0.25/$15 to $2/$2.50/$0.20/$12. Sol unchanged. Cursor's docs state the pool bills third-party API pricing at cost, so the cut flows straight through. Re-confirmed 2026-08-11. |
| Augment Code | Aug 2026 | Luna $1.00/$6.00 to $0.20/$1.20 (80%) and Terra $2.50/$15.00 to $2.00/$12.00 (20%) — the largest single-model rate move since Augment began publishing per-model pricing. The flat $100/month Business plan (up to 50 seats) was unchanged. |
| LiveKit | Aug 2026 | Per-minute denomination, same magnitudes: Luna $0.0040 to $0.0008/min (80%), Terra $0.0101 to $0.0081/min (20%), Sol held at $0.0203/min. The pricing page's default $0.0672/min voice-agent estimate did not move because it runs on Gemma 4 31B, which was not repriced. |
| Vercel | Aug 2026 | v0 Mini fell $1/$5 to $0.20/$1.20 per 1M — landing on Luna's exact new rate — and v0 Pro $3/$15 to $2/$10. Vercel's own note attributes it to the zero-markup AI Gateway passthrough. Because v0's plan cards never show token rates, the repricing is invisible on the pricing page. |
| Netlify | Aug 2026 | Same mechanic, different upstream: Gemini 3.6 Flash halved ($1.50/$7.50 to $0.75/$3.75 per 1M, cache write $0.15 to $0.07) and Gemini 3.7 Flash listed at the identical rate. Netlify converts provider rates at a fixed 180 credits per $1, so the cut passes through to credit consumption without any plan change (Free $0 / Personal $9 / Pro $20 unchanged). |
| Cursor | Aug 2026 | The transmission window collapsed from eleven days to ZERO, and Cursor's last non-pass-through surface converted on the same day. OpenAI cut GPT-5.6 Sol to $4.00/$20.00 on 2026-08-26; Cursor's Other Models pool published $4/$5/$0.4/$20 the same day, and also cut Claude Sonnet 5 from $3/$3.75/$0.3/$15 to $2/$2.5/$0.2/$10 (matching Anthropic's introductory rate) and replaced Gemini 3.6 Flash with 3.7 Flash at roughly half the per-token rate. Separately and more structurally, Cursor's Auto Cost mode DROPPED its flat blended $1.25-in / $6-out rate: all three Auto modes (Cost, Balance, Intelligence) now bill at the routed model's own list price, with the blended rate surviving only as "Legacy Enterprise Auto" and scheduled to sunset 2026-09-07. The corpus's most-cited example of a reseller taking blended margin removed its own blend. |
| Linear | Aug 2026 | A NEW ADOPTER outside the coding-assistant cohort, and the cleanest published statement of the mechanic. Linear replaced Coding Sessions' opaque task-tier estimates ($0.50-$1 for a copy tweak, $3-$5 for a small bug fix, $5+ for complex work) with itemized billing: model tokens at provider-published rates with NO markup, plus a flat $0.25 per 20-minute sandbox block. Code Intelligence also dropped its Beta label. The vendor now charges for the sandbox and the orchestration, and explicitly nothing for the tokens. |
| Composio | Aug 2026 | The markup becomes a NAMED, separately-stated fee rather than a hidden multiple: premium tools moved from a flat 3x multiplier on provider cost to provider cost plus a flat 5% platform fee. Same architecture as zero-markup resale with an explicit thin take-rate on top — the multiplier is disclosed as a percentage instead of being embedded in the rate. |
| Netlify | Aug 2026 | Transmission at catalog scale on a non-OpenAI upstream. Netlify's AI Gateway added OpenRouter on 2026-08-25 (~161 open-weight models at the same fixed 180-credits-per-dollar conversion as Anthropic/Gemini/OpenAI) and repriced dozens of them within one day as OpenRouter moved — DeepSeek V4 Flash 0731 $0.07/$0.18 to $0.03/$0.08, Kimi K2.6 output $2.47 to $3.39, Qwen3-235B $0.23/$2.30 to $0.30/$3.00 — with the catalog growing to ~167 (11 added, 5 removed). Rates moved in BOTH directions in the same pass, which is what a fixed-conversion pass-through looks like when the upstream is a marketplace rather than a single lab. |
| Gumloop | Aug 2026 | The countervailing form: an explicit orchestration take-rate rather than zero markup. Gumloop added a new 8% Orchestration Fee (16% with BYOK) on top of its existing Chat & Reasoning, Tool Call and Compute credits, in the same change that replaced its $37-$1,840 credit slider with a flat $37/mo Pro plan and re-capped overage at 1,000,000 credits ($5,000) per period, off by default. Notably the fee DOUBLES when the customer brings their own key — the opposite of the BYOK discount logic — so the platform prices orchestration higher precisely where it earns no inference margin. |
| OpenRouter | Sep 2026 | Added a Business tier at an 8% platform fee (vs 5.5% Standard) for EU/US in-region routing — the take-rate on pass-through tokens tiered by a governance feature. |
| Netlify | Sep 2026 | AI Gateway cut gpt-5.6-sol to $4.00 / $20.00 per 1M (cache-read $0.50 to $0.40, cache-write $6.25 to $5.00), 12 days after OpenAI's 2026-08-26 cut. |
| Augment Code | Sep 2026 | Cut GPT-5.6 Sol to $4.00 / $20.00 and Claude Sonnet 5 to $2.00 / $10.00 per 1M on its per-token rate card (captured 2026-09-10 vs 2026-08-27). |
| LangChain | Sep 2026 | LLM Gateway sells pay-as-you-go model usage as prepaid Gateway Credits with tokens at face value and a 3.5% processing fee on credit purchases — the margin as an explicit percentage. |
| AssemblyAI | Sep 2026 | LLM Gateway cut GPT-5.6 Luna 80% ($1.00 / $6.00 to $0.20 / $1.20) and Terra 20% ($2.50 / $15.00 to $2.00 / $12.00), matching OpenAI's rates 25 days after OpenAI's change, and grew from 31 to 37 models. |
| Glean | Sep 2026 | Model Hub card (stamped 9/15/2026) cut GPT 5.6 Sol to $4.00 / $20.00 with cache lines in step — about 20 days after the upstream change. |
| Netlify | Sep 2026 | Listed claude-opus-5-5 at $4.00 / $20.00 per 1M one day after Opus 5.5 appeared in Vertex AI's partner catalog. |
Counterexamples
- Perplexity · — — Launched the Gateway API on 2026-08-11 explicitly to STOP passing through: it hosts open-weight models at Perplexity's own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40), inverting its Agent API's zero-markup resale. Three days later (2026-08-14) it also raised Agent API fetch_url 100% back to $0.0005, undoing its own 2026-07-29 cut.
- SambaNova · — — Hosts its own silicon and sets its own rates: on 2026-08-11 it REVERSED its 2026-07-23 gemma-4-31B-it cut, taking the model from $0.22/$0.59 back up to $0.38/$1.15 (+73% input, +95% output) — a move no passthrough vendor could make.
- Sarvam AI · — — Moved upstream rates the opposite direction on its own docs surface: Sarvam-105B repriced roughly 7x higher (Rs 4 / Rs 2.5 / Rs 16 to Rs 29.28 / Rs 10.98 / Rs 73.2 per 1M) on 2026-08-14, while its own marketing pricing page still advertises the old rate.
Trivia
-
OpenAI's $4/$20 GPT-5.6 Sol price is labelled promotional through 2026-11-21. Netlify, Augment Code and Glean all carried it on their cards in September without marking it as dated.
-
AssemblyAI's gateway matched OpenAI's 80% Luna cut 25 days after OpenAI made it (2026-09-20). Cursor had published the Sol cut on the same day OpenAI did.
-
Three companies published the identical two percentage moves on the same calendar day — 2026-08-11 — without coordinating: Augment Code, LiveKit and Vercel all took GPT-5.6 Luna down 80% and Terra down 20%, seven days after Glean did the same, with Cursor's 08-06 move re-confirmed the same day. LiveKit denominates in minutes rather than tokens ($0.0040 to $0.0008/min) and still landed on exactly 80%.
-
Vercel's v0 Mini fell to $0.20/$1.20 per 1M — landing on GPT-5.6 Luna's exact new rate — and the repricing is invisible on Vercel's own pricing page, because v0's plan cards never show token rates at all.
-
The everyday tier launched more expensive than the model it replaced: GPT-5.6 Luna arrived on 2026-07-23 at $1/$6 per 1M against the outgoing GPT-5.4 mini at $0.75/$4.50. The 80% cut twelve days later was a correction to a launch price, and every reseller inherited both moves.
For buyers
Ask which models in a platform's catalog are marked up and which are passed through at cost, and get the answer in writing — it decides whether the vendor can hold your rate when upstream moves, and whether a favourable cut reaches you at all. For a passthrough vendor, monitor the upstream provider's rate card rather than the reseller's: the reseller has no notice obligation for a change it did not make, and the move may not even be visible where you look (Vercel's v0 plan cards show no token rates, so an 80% cut landed with nothing on the pricing page). Then treat the reverse case as the real risk — an at-cost pipe transmits increases with exactly the fidelity it transmitted this cut, and Luna itself launched at $1/$6 above the GPT-5.4 mini it replaced at $0.75/$4.50. Note the expiry you cannot see: OpenAI labels the GPT-5.6 Sol $4/$20 rate promotional through 2026-11-21, and no reseller card marks it as dated. Finally, read the packaging alongside the rate: Glean simultaneously reclassified Luna from Premium to Standard, moving it inside the included 100/user/week allowance instead of always billing FlexCredits, so the effective change was larger than the percentage.
For vendors
Running zero-markup resale means shipping a gateway that rates third-party tokens at cost and a published conversion buyers can audit — Netlify's fixed 180 credits per $1, Cursor's documented at-cost Other Models pool, Vercel's zero-markup AI Gateway. The payoff is that a cheaper upstream reaches your customer automatically with no repricing work; the cost is that you have no model margin and no ability to hold a rate, in either direction. If you take margin instead, you get independence and you should say so on the pricing page: Perplexity launched a Gateway API on 2026-08-11 specifically to stop passing through, hosting open-weight models at its own set rates (kimi-k3 $3.00/$15.00, glm-5.2 $1.40/$4.40) while its Agent API keeps reselling at cost, and SambaNova — on its own silicon — reversed a cut the same day. Either policy is defensible; what isn't is leaving the buyer to guess. The middle path is to put the token line at cost and name your margin as its own line — Composio's provider cost + 5%, Gumloop's 8% Orchestration Fee, LangChain's 3.5% processing fee on credit purchases, OpenRouter's 5.5% Standard and 8% Business fees. And if you pass through, ship a changelog anyway, because your rate card now moves without you.
Outlook — what to watch
Status holds at 432: the falsification test — a zero-markup reseller holding a rate after its upstream moves — is still unmet, and the mechanic has spread past coding assistants (Linear bills provider token rates at no markup plus a $0.25 per 20-minute sandbox fee; Cursor retired its last blended Auto rate). What has changed is the speed: Cursor's same-day pass-through of the Sol cut was the fast end, and 12 to 25 days is now common. The event to watch is 2026-11-21, when OpenAI's promotional Sol rate is labelled to end; if it reverts, this mechanism predicts synchronous increases across resellers. The trend weakens if a zero-markup reseller buffers an upstream move behind a notice period, or if the same synchrony shows up among vendors that take model margin.
Bottom line
The corpus holds two populations that look identical on a rate card: vendors whose model prices are decisions, and vendors whose model prices are consequences. Only the second group moves together — five of them published the same two percentage cuts in August 2026 without making a pricing decision — though the lag ranges from same-day to several weeks, and the margin is reappearing as a named fee rather than vanishing.
FAQ
How can I tell whether a platform marks up model tokens or passes them through at cost?
Look for a stated policy first — Cursor's docs say its Other Models pool bills third-party API pricing at cost, Vercel attributes v0's rates to its zero-markup AI Gateway, and Netlify publishes a fixed 180-credits-per-dollar conversion. Absent a policy, the tells are behavioural: the listed rate matches the provider's public rate to the cent (Vercel's v0 Mini landed on GPT-5.6 Luna's exact new $0.20/$1.20), the change lands the same week as other vendors', and no announcement accompanies it. It matters because a marked-up vendor can hold your rate when upstream moves — and can also keep a cut instead of passing it on.
Why did my AI coding bill drop without any announcement?
Because your vendor probably didn't make a decision. Between 2026-08-04 and 2026-08-11, Glean, Cursor, Augment Code, LiveKit and Vercel each cut GPT-5.6 Luna 80% and GPT-5.6 Terra 20% — a single upstream rate-card move flowing through five at-cost billing pipes. There was nothing to announce, and no notice obligation, because none of the five set those prices.
Does at-cost passthrough transmit price increases too?
Yes, with exactly the same fidelity. GPT-5.6 Luna launched on 2026-07-23 at $1/$6 per 1M — above the outgoing GPT-5.4 mini at $0.75/$4.50 — so every reseller inherited an increase before it inherited the 80% correction twelve days later. Perplexity shows how fast direction can flip: it raised its Agent API fetch_url 100% back to $0.0005 on 2026-08-14, undoing its own 2026-07-29 cut. Zero markup is not a discount guarantee; it is an abdication of rate control in both directions.
If prices pass through, what should I actually monitor?
The upstream provider's rate card, not the reseller's. Then normalise for denomination, because the same move arrives in different units: LiveKit quotes GPT-5.6 Luna per minute ($0.0040 to $0.0008/min) and still lands on exactly 80%, Netlify converts to credits at 180 per dollar, and Vercel's v0 plan cards never display token rates at all. Also check what is not repriced — LiveKit's headline $0.0672/min voice-agent estimate didn't move because it runs on Gemma 4 31B, which wasn't part of the cut.
How quickly do at-cost resellers pass on a model price cut?
Anywhere from the same day to several weeks. Cursor published OpenAI's 2026-08-26 GPT-5.6 Sol cut the same day, while Netlify took 12 days, Augment Code 15 and Glean about 20; AssemblyAI's LLM Gateway carried the earlier Luna and Terra cuts 25 days after OpenAI's change. If a cut matters to your budget, check the reseller's card rather than assuming it has already moved.