Perplexity launches Gateway API for its own hosted open-weight models
Perplexity added a fifth developer API, the Gateway API, offering unified access to open-weight models it hosts itself (DeepSeek, Kimi, GLM) at Perplexity-set per-token prices from $0.13/1M input tokens, distinct from the Agent API's at-cost third-party resale.
Perplexity's developer platform had four API surfaces: Sonar API (token + per-request search fee), Search API ($5.00/1K requests), Agent API (third-party models resold at direct provider rates with no markup, plus metered tool calls), and Embeddings API. No Perplexity-hosted, Perplexity-priced model inference product existed.
A fifth surface, the Gateway API, launched at docs.perplexity.ai/docs/gateway/models: three Perplexity-hosted open-weight models — perplexity/deepseek-v4-flash-0731 ($0.13 input / $0.26 output / $0.028 cache-read per 1M tokens), perplexity/kimi-k3 ($3.00 / $15.00 / $0.30), and perplexity/glm-5.2 ($1.40 / $4.40 / $0.14) — billed per token with no per-request fee, reachable via OpenAI-compatible Chat Completions and Anthropic-compatible Messages endpoints under one API key.
Proof of change
Perplexity AI's pricing pages, as we captured them on two dates.
These values appear in only one of the two captures. That can mean a page-layout difference rather than a price move — read the images, not just the list.
Showing the whole page as captured — scroll either panel, or open it at full size.
- 2 pages had no counterpart in the earlier capture and are not shown.
Both images are our own captures, taken on the dates shown. The pixels are unmodified — nothing is retouched, and where a region is highlighted the marker is drawn over the image, not into it. Open either capture to see it at full resolution.
Perplexity’s developer documentation quietly added a new top-level product on 2026-08-11: the Gateway API, described as giving “unified access to open-weight models hosted by Perplexity through a single endpoint and a single API key.” It sits alongside — not inside — the existing Agent API, Sonar API, Search API, and Embeddings API in the docs sidebar.
The structural distinction matters. Perplexity’s Agent API resells third-party frontier models (OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, NVIDIA) “at direct provider rates with no markup” — Perplexity takes zero model margin there and instead charges for the retrieval/tool layer wrapped around the tokens. The Gateway API inverts that: Perplexity hosts the inference itself for three open-weight models (from DeepSeek, Moonshot AI, and Z.AI) and sets its own per-token rate card, complete with a discounted cache-read rate and reasoning tokens billed at the output rate. It is the first Perplexity developer product priced with real hosting margin rather than at-cost pass-through, positioning Perplexity alongside inference-hosting platforms like Together AI, Fireworks, and DeepInfra for open-weight model serving.
No existing rate moved: Sonar token and request-fee pricing, the $5.00-per-1,000 Search API, all Agent API tool prices, and Embeddings API rates were all re-verified unchanged on the same pass.
Perplexity's developer docs added a fifth API surface, the Gateway API, offering "unified access to open-weight models hosted by Perplexity through a single endpoint and a single API key" via OpenAI-compatible Chat Completions and Anthropic-compatible Messages formats. Unlike the Agent API's zero-markup third-party resale, Gateway models are hosted and priced by Perplexity itself: perplexity/deepseek-v4-flash-0731 ($0.13 input / $0.26 output / $0.028 cache-read per 1M tokens), perplexity/kimi-k3 ($3.00 / $15.00 / $0.30), and perplexity/glm-5.2 ($1.40 / $4.40 / $0.14), billed per token with no per-request fees. No existing Sonar, Search API, Agent API tool, or Embeddings rate moved.