What is it
Per-message pricing is a billing unit where each individual message or reply in a conversation is metered, common in AI chat and voice platforms.
The message is the most intuitive unit in conversational AI. It mirrors how users already think — “I sent a message, I got a reply” — and maps onto the single LLM call that drives inference cost. Where token pricing is precise but opaque to non-technical buyers, message pricing trades granularity for a meter that is human-readable before the buyer opens the docs.
The unit shows up in two commercial contexts. Consumer AI platforms — Poe, Grok, Janitor AI, and the now-sunset Cognosys — use message caps to differentiate subscription tiers rather than billing per message, so the message rate is the invisible floor even when no dollar-per-message line item ever appears. Voice and chat API platforms take the opposite tack: Retell AI and Vapi publish explicit per-message rates for their chat-agent products — $0.002 and $0.005 respectively — treating the message as a pure usage meter alongside per-minute voice billing.
Edge cases reveal how elastic “message” can be. Intercom Fin carries messages through a Proactive Support Plus add-on, separate from its headline per-resolution model, and Pi rate-limits on messages purely to control cost — its effective rate is $0, and the meter never converts to a charge. A message meter in the frontmatter does not imply a per-message bill. For how to pick a unit like this, see choosing the right usage metric.
How it works
A message meter counts conversational turns and multiplies by a rate. The rate can be an explicit dollar figure (API platforms), a points cost that varies by model (Poe), or an implicit ceiling encoded as a tier cap (consumer apps). What differs across vendors is what actually gets counted and how the price is exposed.
| Dimension | What it controls | Example on this page |
|---|---|---|
| Rate exposure | Whether the message price is a public dollar figure or hidden inside a subscription | Retell AI publishes “from $0.002/AI msg”; Grok hides it behind flat tiers |
| Unit of counting | User turn, AI reply, or both | Retell counts each AI message; Poe charges per message to a bot |
| Rate variability | Flat per message vs. model-dependent | Vapi is a flat $0.005/msg; Poe’s points cost swings ~10–20 pts (light text) to thousands (frontier/video) |
| Packaging | Metered directly vs. bundled into a cap | Retell/Vapi meter directly; Cognosys capped Free at 100 msg/mo, Pro (~$15/mo) at ~1,000 |
The clearest direct-metered example is Retell AI. Its chat agents are billed per AI message from $0.002, and the effective rate depends on the model — a GPT 5.1 chat message computes to about $0.013. There is no platform fee, no contract, and new accounts get $10 in free credits.
Unit math (Retell chat agent): 5,000 GPT 5.1 messages × $0.013/msg ≈ $65. The same volume at the $0.002 floor rate (a lighter model) is $10 — the model choice is the dominant cost lever, not the message count.
Poe shows the points-abstraction variant: one subscription buys a monthly (or daily) compute-points allowance, and every message spends points at a published, model-dependent rate. A lightweight text model costs ~10–20 points per message, so a $19.99 plan stretches to tens of thousands of cheap messages — but a single high-end video generation can cost more points than thousands of text messages. The message is the meter; the model mix is the real cost driver.
Companies using this
Eight companies in the corpus list messages as a billing unit, split between voice/chat API platforms that meter it directly (Retell AI, Vapi) and consumer apps that use message caps to gate subscription tiers (Poe, Grok, Janitor AI, Cognosys).
Patterns observed
-
Direct metering lives in the API layer, caps live in the app layer. The only companies quoting a public per-message dollar rate — Retell AI and Vapi — are developer platforms selling chat agents. Every consumer product expresses the unit as a tier cap instead: Cognosys at 100 messages/month, Janitor AI at ~50/day, Grok at ~10 prompts per 2 hours.
-
The message is almost never the only meter. On every company here,
messagesrides alongside another unit — minutes for Retell AI and Vapi, points for Poe, tokens for Grok and Janitor AI, resolutions for Intercom Fin. Message pricing is a legibility layer over an underlying token or compute cost, not a standalone model. -
Model choice, not message volume, drives the bill. Where the rate is model-dependent — Poe’s points, Retell AI’s per-model chat rates — the same message count can cost 5–10× on the model selected. Buyers budgeting on message count alone consistently underestimate frontier-model spend.
Counterexamples & variants
The clearest counterexample is Pi. It lists messages as a billing unit, but the app has never been monetized — Inflection AI rations capacity with rate limits and cooldowns rather than dollars, so the effective rate is $0 and the meter never converts to a bill. It is a pure cost-control meter, not a pricing model.
Cognosys is a lifecycle variant: it ran a clean message-metered freemium ladder (Pro ~1,000 msg/mo at $15, Ultimate unlimited at $59) but is now sunset after the Cohere acquisition, so its rate card is historical rather than purchasable. Intercom Fin inverts the emphasis: its economics run on $0.99-per-resolution outcome pricing, with messages surfacing only in a peripheral Proactive Support add-on (500/mo at $99). The message unit is real but commercially minor.
What this means for buyers vs vendors
For buyers
Confirm exactly what a “message” is before you model spend — Retell AI counts each AI reply, Poe charges per message to a bot, and consumer caps count prompts. Then model on your worst-case model mix, not message volume, since a model-dependent rate is what actually moves the bill. Use the pricing calculator to stress-test a high-volume month and read usage invoicing and billing cycles to understand how overages and caps are reconciled.
For vendors
Per-message billing fits when your unit cost is predictable per exchange and your buyer is non-technical enough that tokens would confuse them — the message is a legibility win that Retell AI turned into a self-serve growth lever. Pair it with a second meter that tracks real cost (minutes, points, or tokens) so a frontier-model exchange doesn’t erode margin, and start with the fundamentals in introduction to usage-based pricing. Expect to publish caps or free-tier limits, since every consumer app on this page gates messages rather than charging for them directly.
| Company | Product | Pricing model | Billing units | Free tier | Verified |
|---|---|---|---|---|---|
| ActiveCampaign | Marketing automation, email marketing, and sales CRM platform priced by contact count | No | 2026-07-12 | ||
| Bland AI | AI phone call automation platform — inbound and outbound voice agents at scale | Yes | 2026-07-21 | ||
| Braze | Enterprise customer-engagement platform for cross-channel messaging (email, push, SMS, in-app), with BrazeAI (Sage AI) bundled. | No | 2026-07-12 | ||
| Cognosys | Autonomous AI agents (rebranded Ottogrid, acquired by Cohere) | Yes | 2026-07-30 | ||
| Constant Contact | Email and digital marketing platform for small businesses, with a separate Lead Gen & CRM product (ex-SharpSpring) and a bundled AI content assistant | No | 2026-08-06 | ||
| Customer.io | Customer engagement platform combining Journeys (behavioral messaging), an AI Agent, and Data Pipelines (CDP) | No | 2026-07-12 | ||
| Grok | xAI's consumer and business AI assistant | Yes | 2026-06-16 | ||
| Intercom Fin | Fin AI Agent for customer service | No | 2026-08-04 | ||
| Iterable | Cross-channel marketing automation and AI customer engagement platform (email, SMS, push, in-app, web, OTT) | No | 2026-07-12 | ||
| Janitor AI | Consumer AI character chat / roleplay platform | Yes | 2026-06-16 | ||
| Keap | All-in-one CRM, sales, and marketing-automation platform for small businesses | No | 2026-07-06 | ||
| Klaviyo | B2C marketing CRM for email, SMS, and push, priced on active profiles plus channel volume, with bundled predictive analytics and Klaviyo AI. | Yes | 2026-07-12 | ||
| Microsoft Dynamics 365 | Microsoft's enterprise CRM + ERP suite — Sales, Customer Service, Field Service, Business Central, Finance and Supply Chain, with Copilot woven in | No | 2026-07-06 | ||
| Pi | Pi — personal, emotionally intelligent AI assistant (consumer app) | Yes | 2026-06-16 | ||
| Poe | Multi-model AI chat subscription (by Quora) | Yes | 2026-07-28 | ||
| Retell AI | Conversational voice-agent API platform | No | 2026-07-22 | ||
| Snowflake Cortex | AI functions and model APIs on Snowflake | Yes | 2026-07-06 | ||
| Vapi | Voice AI infrastructure for developers | No | 2026-06-09 |
Explore this theme in the knowledge graph
FAQ
What is per-message pricing in AI products?
Per-message pricing charges each individual message or reply in a conversation as the billable unit. Consumer AI apps like Poe and Grok use message caps to gate subscription tiers, while voice AI platforms like Retell AI and Vapi charge fractional cents per AI message on their chat agents.
How does per-message pricing differ from per-token pricing?
Tokens measure the raw units the underlying model processes; a message is a complete conversational turn — a prompt plus its reply. One message can consume a few hundred to several thousand tokens depending on context length, so per-message rates are higher in absolute terms but far more legible to non-technical buyers.
Which AI companies charge per message?
Retell AI charges from $0.002 per AI message on chat agents; Vapi charges $0.005 per SMS/chat message; Poe uses compute points where each message to a model costs a published point rate. Consumer apps like Poe, Grok, Janitor AI, and the sunset Cognosys bundle messages into subscription tiers rather than billing per message directly.
How much does per-message chat cost on voice AI platforms?
Retell AI bills chat agents from $0.002 per AI message (GPT 5.1 works out to about $0.013 per message), and Vapi charges a flat $0.005 per SMS or chat message on top of its $0.05-per-minute voice hosting. Both give new accounts $10 in free credits to start.
What are the risks of per-message pricing for buyers?
The main risks are unpredictable costs when conversation volume spikes and bill shock when a single session with a frontier model burns through a points budget. Buyers should confirm whether a message counts each user turn, each AI reply, or both — Poe, Retell, and Cognosys all define it differently.
Related billing units
- Credit-Based BillingA billing unit where customers pre-purchase or are allocated a pool of credits that deplete as they use the product, often at variable rates per feature.
- Token-Based PricingA billing unit common in LLM and AI products, where customers are charged per input and output token processed.
- Per-Seat PricingA billing unit where the vendor charges a fixed fee per named user, regardless of how much each user consumes.
- Per-Resolution PricingA billing unit unique to AI customer-support products, where the vendor charges only when an AI agent resolves a customer issue without escalation.
- Bandwidth-Based PricingA billing unit where customers are charged per gigabyte of data transferred out of the platform.
- Per-Function-Invocation PricingA billing unit where customers are charged per serverless function invocation, often combined with a separate compute-time charge.
- CPU-Hour PricingA billing unit where customers are charged for the CPU time their workloads consume, typically measured in vCPU-seconds or vCPU-hours.
- GB-Hour PricingA billing unit where customers are charged for the memory their workloads consume over time, measured in gigabyte-hours.
- GPU-Hour PricingA billing unit where customers are charged for GPU time consumed, typically measured per-second or per-hour by GPU type.
- Per-API-Call PricingA billing unit where customers are charged per API request, regardless of payload size or processing time.
- Per-GB Storage PricingA billing unit where customers are charged per gigabyte of data stored on the platform per month.
- Media-Minute PricingA billing unit where customers are charged per minute of audio or video processed — used by speech, voice, and video AI vendors.
- Per-Request PricingA billing unit where customers are charged per request served — the generic meter for inference endpoints, search, scraping, and browser infrastructure.
- Per-Event PricingA billing unit where customers are charged per event ingested — the native meter of observability and billing-infrastructure platforms.
- Vector Storage PricingA billing unit where customers are charged for vectors stored or indexed — the storage dimension of vector database pricing.
- Per-Character PricingA billing unit where customers are charged per character of text processed — the standard meter for text-to-speech and translation.
- Per-Document PricingA billing unit where customers are charged per document processed or generated — common in AI writing, SEO, and document-intelligence tools.
- Per-Page PricingA billing unit where customers are charged per page crawled, parsed, or rendered — the meter for web scraping and document parsing.
- Per-Transaction PricingA billing unit where customers are charged per financial or billing transaction processed — the meter of billing and accounting platforms.
- Active-User PricingA billing unit where customers are charged per monthly or daily active user rather than per provisioned seat.
- Per-Task PricingA billing unit where customers are charged per task an automation or agent executes — Zapier's historical unit, now spreading to AI agents.
- Per-Unit PricingA billing unit used by robotics, hardware AI, and some SaaS companies where the metered object is a physical or abstract 'unit' — a robot deployed, a device sold, or a defined deliverable.
- Workflow Execution PricingA billing unit where each end-to-end workflow or automation run is metered and billed, regardless of the compute steps it contains.
- Per-Invoice PricingA billing unit used by billing infrastructure platforms where each invoice generated or processed is metered as the primary cost driver.
- Per-Action PricingA billing unit where each discrete action taken by an AI agent or automation is metered — common in browser automation and agentic workflow tools.
- Per-Image PricingA billing unit where each AI-generated image is metered, common in image generation APIs and multimodal AI platforms.
- Per-Conversation PricingA billing unit where each complete customer conversation — from first message to resolution — is metered as a single chargeable event.
- Per-Record PricingA billing unit where each data record processed, labeled, or extracted is metered — common in data platforms and web scraping services.
- Per-Word PricingA billing unit common in translation and localization platforms where the metered object is the word count of content processed.
- Per-Video PricingA billing unit where each AI-generated video is metered, common in video generation and synthetic media platforms.
- Milestone-Based PricingA billing unit used in drug discovery and biotech AI where payment is tied to achieving defined research milestones rather than time or compute consumed.
- Per-Outcome PricingA billing unit where payment is triggered by verified outcomes delivered — distinct from outcome-based pricing models, this refers specifically to 'outcomes' as a countable billing unit.
- Per-Datapoint PricingA billing unit where each individual data measurement or signal ingested is metered — common in cloud cost intelligence and ML evaluation platforms.
- Per-Interaction PricingA billing unit where each patient-agent or user-agent interaction is metered, common in healthcare AI and customer engagement platforms.
- Data Licensing PricingA pricing structure where access to proprietary datasets or data assets is licensed separately from the software or services, common in AI training data and clinical data platforms.
- Robot-Hour PricingA billing unit where each hour a robot or autonomous system operates is metered — the robotics equivalent of a GPU-hour.
- Per-Contact PricingA billing unit where each contact or lead in the database is metered, common in AI sales development and outbound automation platforms.
- Per-Mailbox PricingA billing unit where each connected email mailbox or sending account is metered, common in AI outbound sales and email automation platforms.
- Browser-Hour PricingA billing unit where each hour of headless browser compute time is metered, common in web scraping and browser automation platforms.
- Per-Generation PricingA billing unit where each AI-generated creative asset — image, video, or design — is counted as a 'generation' and metered accordingly.
- Per-Ticket PricingA billing unit where each customer support ticket handled by an AI agent is metered — common in AI customer service platforms.
- Per-Log PricingA billing unit where each LLM request log ingested or stored is metered — common in AI observability and evaluation platforms.
- Per-Trace PricingA billing unit where each distributed trace — a complete record of an LLM request chain — is metered, common in AI observability platforms.
- Per-IP PricingA billing unit where each IP address or proxy endpoint allocated is metered — used by web scraping proxy providers.
- Per-Device PricingA billing unit where each hardware device or endpoint connected to the AI platform is metered.
- Per-Case PricingA billing unit used in legal AI platforms where each case or matter processed by the AI is metered.
- Per-Report PricingA billing unit where each AI-generated report or analysis document is metered as a discrete output.