Perplexity adds a Gateway API for self-hosted open-weight models
Perplexity quietly launched a Gateway API hosting open-weight models (Moonshot AI's Kimi K3, Z.AI's GLM 5.2) at its own per-token rates, a fifth metered developer surface distinct from the Agent API's at-cost third-party resale.
Four metered developer surfaces: Sonar API, Search API, Agent API (zero-markup third-party model resale), Embeddings API.
Five metered developer surfaces: adds Gateway API — perplexity/kimi-k3 ($3.00 input / $15.00 output / $0.30 cache-read per 1M tokens) and perplexity/glm-5.2 ($1.40 input / $4.40 output / $0.14 cache-read per 1M tokens) — billed at Perplexity's own published rates rather than an at-cost passthrough.
Perplexity’s developer docs added a new Gateway API product line sometime between the 2026-07-29 and 2026-08-05 checks — a per-token rate card for open-weight models Perplexity hosts itself, reachable through both Chat Completions and Messages endpoints under a shared creator/model-name id. Unlike the Agent API’s zero-markup resale of third-party frontier models, Gateway API pricing is Perplexity’s own to set: the initial catalog covers Moonshot AI’s Kimi K3 and Z.AI’s GLM 5.2, with a live GET /models endpoint that doubles as the request allowlist.
Perplexity's developer docs added a fifth API surface, the Gateway API, offering "unified access to open-weight models hosted by Perplexity through a single endpoint and a single API key" via OpenAI-compatible Chat Completions and Anthropic-compatible Messages formats. Unlike the Agent API's zero-markup third-party resale, Gateway models are hosted and priced by Perplexity itself: perplexity/deepseek-v4-flash-0731 ($0.13 input / $0.26 output / $0.028 cache-read per 1M tokens), perplexity/kimi-k3 ($3.00 / $15.00 / $0.30), and perplexity/glm-5.2 ($1.40 / $4.40 / $0.14), billed per token with no per-request fees. No existing Sonar, Search API, Agent API tool, or Embeddings rate moved.