Ask
Sharpens 13 companies · First observed June 2025 · Updated September 2026 Explore in the graph

BYO-API keys zero out the model-cost meter

Quick answer

A cohort of agentic and workflow platforms now runs two meters: an orchestration credit they keep, and a passthrough model/data cost they mark up. The tell is that Bring-Your-Own-Key zeroes or halves only the passthrough meter — six corpus vendors expose this lever, while first-party APIs (OpenAI) have none and some seat+credit tools (Cursor) are pulling it back.

6 vendors let buyers zero or halve the model meter via BYOK

What's happening — and why

What's happening: platforms that resell model inference are increasingly splitting the bill into two separate meters. One meter charges for their own orchestration — Clay's Actions, Relevance AI's Platform Credits, Gumloop's workflow nodes — and a second meter charges for the model and data they pass through (Clay's Data Credits, Relevance's Vendor Credits). The diagnostic is what happens when you bring your own LLM API key: BYOK zeroes or halves only the passthrough meter, never the orchestration one. Clay's BYOK eliminates Data Credit cost while still burning Actions; Relevance's BYOK zeroes Vendor Credits; Gumloop's BYOK cuts agent AI-model credits by exactly 50% (its native nodes already cost 0 credits). Vectara and Byword extend the same logic to bundled models and writing.

Why: the split makes the vendor's real value metric legible. By offering a key that zeros the model meter, the platform concedes it is not the cost center on inference — it is charging for the plumbing, not the tokens. The lever isn't new (Byword sold an 'Unlimited' plan on the customer's own GPT-4 keys at $2,499/mo back in December 2023); what changed is that mid-market tools now expose it openly instead of pricing it as a five-figure enterprise escape hatch.

How it works

your usage orchestration Actions / nodes passthrough model / data always billed BYOK zero / -50% vendor bill orchestration only model provider you pay wholesale
Two meters: BYOK zeroes (or halves) the model passthrough while orchestration credits keep billing — you pay the model provider directly.

Evidence over time

14 supporting · 7 counter — hover or tap a point for detail, click to jump to the row.

supports ↑ challenges ↓ 2025 2026
supporting evidence counterexample

Evidence

Company Date What happened
Clay Jun 2026 Two-meter model (Actions + Data Credits) where bringing your own API keys 'eliminates Data Credit cost while still consuming Actions' — BYOK zeros the data/AI passthrough meter but not the orchestration meter.
Relevance AI Jun 2026 Bills Platform Credits + Vendor Credits (raw model cost passed through at wholesale); bringing your own LLM keys on any paid plan 'zeroes this out' — buyers can opt out of one of the two metered dimensions entirely.
Gumloop Jun 2026 Bringing your own API key cuts agent AI-model credit costs by 50%; native workflow nodes (logic, loops, Sheets, Slack) already cost 0 credits — BYOK halves the model-driven portion of the bill.
Byword Jun 2026 Credit-priced rebuild offers a BYO-API path (Claude + Gemini) on an Unlimited plan; the BYO-key idea is long-standing here — a Dec-2023 archive already sold an 'Unlimited' plan running on the customer's own GPT-4 keys at $2,499/mo.
Vectara Jun 2026 Every tier bundles Vectara's own retrieval + generative LLMs but offers Bring Your Own Model (ChatGPT, Claude, Gemini) — letting enterprises substitute their own model contract for the bundled one.
Lemlist Jun 2026 Monetizes the data you pull through credits rather than the bundled 650M+ database; BYO-data/key economics let buyers collapse the 'data tool + sender' stack and avoid double-paying for model/data passthrough.
Tabnine Jun 2026 LLM token usage is unlimited when you bring your own model (on-prem or your own cloud endpoint); only Tabnine-provided LLM access is metered — at the provider price plus a 5% handling fee. BYOK zeroes the model meter on a per-seat coding platform.
Deepgram May 2026 Voice Agent BYO tiers price the same split per minute: Standard – BYO TTS at $0.065/min vs $0.075/min fully managed — bring your own component, pay less per minute.
Synthflow Jun 2025 Bringing your own Twilio drops the telephony meter to $0.00/min on an unbundled per-minute stack — BYO zeroes one of three meters.
Aider Jun 2026 Open-source CLI AI pair programmer (Apache 2.0) with zero platform fee — the only cost is the user's own BYOK tokens paid directly to providers (OpenAI, Anthropic, Google, DeepSeek). The most extreme form of BYOK: there is no platform/orchestration meter at all, just passthrough.
Continue Jun 2026 Team plan supports BYOK on Company tier with SAML/SSO — the open-source coding assistant lets enterprises supply their own model keys to zero the per-token passthrough.
OpenRouter Jul 2026 Reworked its bring-your-own-key free allotment from a request count (1M req/mo PAYG, 5M Enterprise) to a dollar cap ($25,000/mo of list-price inference PAYG, $200,000 Enterprise), then a 5% fee above it. The take-rate on BYOK now scales with inference value, not request volume — the router keeps a thin percentage on model cost it does not supply, matching the fee to passthrough economics. (Provider count also ticked 60+ → 70+.)
Bland AI Jul 2026 New adopter of the BYO-component mechanic, on telephony rather than the model. Bland's per-minute rate previously advertised as all-in — 'Telephony — PSTN and SIP supported' listed under 'Everything included in your per-minute rate' — now covers LLM ('no token charges'), STT and TTS only, with the pricing-page FAQ rewritten to read that the rate 'covers the LLM, STT, and TTS in one number' and that 'Telephony is billed separately, on your own carrier or Bland's at pass-through cost.' No plan price moved (Start free at $0.14/min, Build $299/mo platform fee at $0.12/min, Scale $499/mo at $0.11/min, Enterprise contracted to volume), so the bundle shrank at a constant price — the same split Synthflow's BYO-Twilio established, now with a second adopter.
Gumloop Jul 2026 Lever intact, on-ramp closed: the 50% BYOK discount on agent AI-model credits is unchanged and the Pro credit slider held at $37 (20,000 credits) to $1,840 (1M credits) across four capture cycles, but the permanent $0/month, 5,000-credit Free plan was discontinued in favour of a 14-day Pro trial offered only on monthly billing. Because Gumloop's native workflow nodes already cost 0 credits, the credit meter only ever measured model passthrough — so a buyer wanting to test the BYOK halving now has to reach a paid plan first.

Counterexamples

  • Cursor (Anysphere) · Aug 2026 — This trend's BYOK-withdrawal counterexample reached the same endpoint without the lever. Cursor's Models & Pricing docs state that the Other Models usage pool passes third-party model API pricing through AT COST, and the window proved it: GPT-5.6 Luna fell about 80% ($1/$1.25/$0.10/$6 to $0.20/$0.25/$0.02/$1.20 per 1M input/cache-write/cache-read/output) and GPT-5.6 Terra about 20% ($2.50/$3.125/$0.25/$15 to $2/$2.50/$0.20/$12), with every other model in both pools and every plan price unchanged. Augment Code, LiveKit Inference and Glean published the identical two percentages on the same two models within the same five days — the signature of four independent vendors carrying no markup on that dimension. A buyer at Cursor cannot bring a key and cannot save anything by doing so, because there is nothing to strip. Withdrawing BYOK and conceding wholesale produce the same margin disclosure; only the second one is visible on the rate card.
  • Perplexity AI · Aug 2026 — The wholesale concession turns out to be surface-specific. Three weeks after widening Agent API third-party resale to 7 providers 'at direct provider rates with no markup', Perplexity launched a fifth developer surface — the Gateway API — for open-weight models it hosts itself, billed at Perplexity's OWN published per-token rates rather than as an at-cost passthrough: perplexity/deepseek-v4-flash-0731 at $0.13 input / $0.26 output / $0.028 cache-read per 1M, perplexity/kimi-k3 at $3.00/$15.00/$0.30, perplexity/glm-5.2 at $1.40/$4.40/$0.14, reachable through OpenAI-compatible Chat Completions and Anthropic-compatible Messages endpoints under one key. By 2026-08-14 the catalog had grown to five with two NVIDIA additions (nemotron-3.5-lightning-30b-a3b at $0.0115/$0.17/$0.00115, undercutting everything else in the catalog, and nemotron-3-ultra-550b-a55b at $0.25/$2.50), the three original prices unchanged. Perplexity now runs zero-markup resale and margin-bearing hosting on the same developer platform, which bounds this trend's counterexample: 'the vendor charges wholesale' is a statement about a product line, not a company.
  • OpenAI · May 2026 — First-party model API — there is no 'bring your own key' lever; you pay per token directly, with no orchestration/passthrough split to opt out of.
  • Cursor (Anysphere) · Feb 2026 — Coding tool restricts/deprecated raw BYOK on its managed plans in favor of bundled credits — the trend toward letting buyers zero the model meter is not universal; some seat+credit tools pull the lever back.
  • Lovable · Jul 2026 — Collapsed the two-meter split the lever depends on. Previously a workspace held credits for agent messages plus a SEPARATE dollar balance for Lovable Cloud (hosting, database, storage, network, compute, realtime) and for the AI gateway that deployed apps call — the passthrough half. Lovable converted remaining dollar balances into credits at the plan's credit rate and now issues the monthly Cloud and AI allowances as credits (20 Cloud + 4 AI per month on Free, Pro and Business) on top of 5 daily build credits, with every top-up control credit-based. Lovable states the underlying cost of running projects has not changed and calls the Cloud/AI grants 'a temporary offering' — but with one currency across orchestration and passthrough, there is no longer a passthrough meter for a key to zero.
  • Perplexity AI · Jul 2026 — Makes the lever economically pointless by conceding wholesale outright: Agent API third-party model resale widened from 4 providers (OpenAI, Anthropic, Google, xAI) to 7, adding Z.AI, Moonshot AI and NVIDIA, all 'at direct provider rates with no markup'. BYOK exists to strip a passthrough markup; where the vendor publishes zero markup, bringing a key saves nothing. The margin transparency this trend treats as BYOK's payoff is reached here without the lever — which bounds the trend to platforms that DO mark up passthrough.
  • Braintrust · Jul 2026 — Removed the incentive to bring a key, one day before its own deadline. GLM-5.2 — an open-source reasoning model offered without requiring a customer's own API key — was documented in both the billing FAQ and the Plans-and-limits page as available 'through July 31, 2026', with no continuation promised. The 2026-07-30 capture shows that clause deleted from both surfaces; the FAQ now says GLM-5.2 'is available as a built-in model under the Braintrust provider' with no end date. The pool funding it was renamed from 'Topics credit' to a shared 'Model credits', and the pricing page swapped its explicit per-tier token-rate line for a generic 'then token rates' plus a details link. Free built-in model access competes directly with BYOK.

Trivia

  • Tabnine (verified 2026-06-08) prices the no-BYO path with unusual precision: bring your own model and LLM usage is unlimited; use Tabnine-provided model access and you pay the provider's actual price plus exactly 5% handling — one of the only vendors to publish its passthrough markup as a number rather than hide it inside a credit.

  • Clay's BYOK (2026-06-02) zeroes the Data Credits meter but still burns Actions — the single sharpest proof in the corpus that these vendors charge for orchestration, not inference: a buyer who supplies their own model key pays Clay nothing for the AI/data and everything for the plumbing, exactly inverting the assumption that the model is the cost center.

  • Gumloop (2026-06-02) puts a precise number on the lever almost no other vendor discloses: bringing your own API key cuts agent AI-model credits by exactly 50%, and its native workflow nodes (logic, loops, Sheets, Slack) already cost 0 credits — so the credit meter only ever measured the model passthrough, which BYOK halves.

  • Byword shows the lever is six years old, not a 2026 invention: a December-2023 archive already sold an "Unlimited" plan that ran on the customer's own GPT-4 keys at $2,499/mo — meaning BYO-key economics predate the agentic-platform wave that made them mainstream, and the only thing that changed is that mid-market tools (Clay, Relevance, Gumloop) now expose it instead of pricing it as a five-figure enterprise escape hatch.

  • The threat to this lever changed shape in July 2026. It is no longer that vendors withdraw BYOK (Cursor's route) but that they stop publishing a passthrough meter for BYOK to zero. Lovable (2026-07-21) merged its dollar-denominated AI-gateway and Cloud balances into its single credit balance, converting the remaining dollars at the plan's credit rate — so orchestration and passthrough now share one currency and a buyer cannot see which half a key would offset. A lever needs two meters to pull between.

  • Perplexity (2026-07-21) reached this trend's endpoint without the lever: it widened Agent API third-party model resale from 4 providers to 7 — adding Z.AI, Moonshot AI and NVIDIA to OpenAI, Anthropic, Google and xAI — explicitly "at direct provider rates with no markup." Where the vendor already passes model cost through at wholesale, bringing your own key saves exactly nothing. Zero-markup resale is the same margin disclosure BYOK forces, arrived at from the opposite direction.

  • Bland AI (2026-07-21) is the corpus's second BYO-carrier adopter after Synthflow, and the cleanest example of a bundle shrinking without a price moving. Every plan price held — Start free at $0.14/min, Build $299/mo platform fee at $0.12/min, Scale $499/mo at $0.11/min — but the "everything included in your per-minute rate" strip dropped its fourth item. Telephony is now "billed separately, on your own carrier or Bland's at pass-through cost", so the same per-minute number buys LLM, STT and TTS only.

  • Braintrust (2026-07-30) removed the reason to bring a key at all — one day before its own deadline. As of the 2026-07-22 capture both the billing FAQ and the Plans-and-limits docs said GLM-5.2 (offered without requiring a customer's own API key) was available "through July 31, 2026", with no continuation promised. On 2026-07-30 that clause is gone from both surfaces and the FAQ simply says GLM-5.2 "is available as a built-in model under the Braintrust provider," full stop — while the pool funding it was renamed from "Topics credit" to "Model credits."

See all pricing trivia

For buyers

If a platform exposes a BYOK lever, price the workload both ways. Heavy-inference jobs usually win by bringing your own key and paying the model provider at wholesale — you keep only the orchestration meter (Gumloop cuts model credits 50%; Clay and Relevance zero theirs entirely). The trap is platforms that mark up passthrough with no BYOK escape hatch; there the 'credit' hides a model markup you can't opt out of. Ask explicitly whether BYOK exists and which meter it touches before you size a plan.

For vendors

Running this play needs a clean two-meter architecture — an orchestration credit you keep, plus a passthrough meter buyers can route around with their own API key. Be explicit about which meter BYOK affects (Clay: zeros Data Credits, keeps Actions; Gumloop: halves model credits, native nodes already free) so the value story stays legible. The strategic risk is the Cursor counter-move: if your differentiation is bundled-credit margin on inference, exposing BYOK concedes you aren't the cost center and invites buyers to unbundle — weigh transparency against the markup you're protecting.

Outlook — what to watch

Expect more mid-market agentic and workflow tools to expose BYOK as a transparency and margin signal, normalizing the two-meter split. The status flips from new to holds if more corpus vendors ship an explicit passthrough-zeroing key, or to weakens if seat+credit tools follow Cursor and pull raw BYOK back to defend bundled-credit revenue — betting buyers prefer one bill over a cheaper assembled one. Watch whether any vendor publishes its dollar-to-credit conversion alongside the BYOK option, which would make the orchestration markup fully legible.

Bottom line

Six corpus vendors (Clay, Relevance AI, Gumloop, Byword, Vectara, Lemlist) split the bill into orchestration vs passthrough meters and let buyers zero or halve the model passthrough with their own API key. It's a transparency and margin signal — the platform admits it charges for plumbing, not tokens — but it isn't universal: first-party APIs have no such lever, and some seat+credit tools are pulling it back.

FAQ

What does 'bring your own key' (BYOK) do to my bill?

On platforms that run two meters, BYOK zeroes or halves the model/data passthrough meter while leaving the orchestration meter untouched. Clay's BYOK eliminates Data Credit cost but still burns Actions; Relevance AI's BYOK zeroes Vendor Credits; Gumloop's BYOK cuts agent AI-model credits by exactly 50%.

Why would a platform let me zero out its model cost?

Because the model was never their cost center. Offering a key that zeros the passthrough meter makes their real value metric explicit — they charge for orchestration (the plumbing), not the inference. It's a transparency and margin signal: you pay them for the workflow and pay the model provider directly at wholesale.

Do all AI platforms offer BYOK?

No. First-party model APIs like OpenAI have no BYOK lever by definition — you pay per token directly with no passthrough split to opt out of. And the direction is contested: some seat+credit tools, like Cursor, have restricted or deprecated raw BYOK to defend bundled-credit revenue.

Is BYOK pricing a new 2026 invention?

No — it's at least six years old. Byword's December-2023 archive already sold an 'Unlimited' plan running on the customer's own GPT-4 keys at $2,499/mo. What changed in 2026 is that mid-market tools (Clay, Relevance, Gumloop) now expose the lever openly instead of pricing it as a five-figure enterprise escape hatch.

All trends