A partner running a white-label AI voice service for medical clinics was 60 minutes into evaluating a vendor when the conversation broke. The vendor priced per minute. The partner sold flat-rate to his clinics. There was no path to a contract. "I can't buy unpredictable and sell unpredictable," he said. The vendor lost the deal. The buyer lost three weeks. The category lost a customer who was ready to buy.
This pattern is the most common pricing-model mismatch in AI voice today. Vendors price for their margin reality. Buyers want a different reality. Most evaluations fail before the pricing-model question gets asked, and most that get asked end with someone walking away.
This post is the practical breakdown. Why the AI voice category started per-minute, what per-minute breaks for buyers, what flat-rate breaks for vendors, and the hybrid models that actually work for both sides.
Why AI voice priced itself by the minute
The AI voice category inherited per-minute pricing from its component parts. Three layers of cost are mostly per-minute or per-token:
- LLM inference. Charged per input + output token. A 4-minute conversation generates roughly 2,000 to 3,500 tokens. At current Flash-tier frontier-model rates (roughly $0.30 per million input tokens, $2.50 per million output), inference runs $0.005 to $0.02 per call.
- Speech-to-text. Charged per audio minute. Major vendors run $0.006 to $0.024 per minute depending on model tier and language.
- Text-to-speech. Charged per character or per minute of synthesized audio. ElevenLabs and similar premium voice tiers run $0.05 to $0.30 per minute synthesized.
Stack these three layers, add telephony (Twilio sits around $0.008 to $0.015 per minute) and infrastructure overhead, and you land at $0.10 to $0.50 per minute in raw vendor cost depending on voice tier and STT model. The early platforms (Vapi, Bland, Synthflow, Retell) priced at the floor of this range, marking up 2 to 4x on infrastructure. This is pure infrastructure pricing where the buyer brings their own prompt, model, and integration.
It works for developers building applications. It does not work for businesses buying receptionists. The two groups have different needs.
What per-minute pricing breaks for buyers
The mismatch lives in the buyer's planning model. A small business buys an answering service to replace a receptionist. The mental frame is: "I have a receptionist budget of $500 a month, I get a receptionist." Per-minute pricing punishes the buyer for traffic the buyer cannot fully predict.
Three specific scenarios where per-minute breaks:
- Volatile traffic. A real estate brokerage gets 200 calls in March (spring buying season) and 80 in October. Per-minute pricing makes the bill volatile. The operator cannot budget. CFO unhappy.
- White-label resale. A partner reselling AI voice to multiple end-customers needs predictable cost-per-customer to set predictable resale prices. The Affilia conversation cited above is the canonical version of this. Per-minute breaks the resale economics outright.
- Heavy first-month use. A new buyer wants to test heavily in the first week (script tuning, scenario testing). Per-minute pricing punishes the evaluation phase, making test runs feel like burning money.
The pattern in each: the buyer wants cost predictability. Per-minute trades predictability for vendor margin protection. The vendor wins on the pricing argument; the buyer walks.
What flat-rate pricing breaks for vendors
Flat-rate is not the answer. It breaks the other side of the same equation.
Three things flat-rate breaks:
- Margin on heavy users. A flat $200 a month plan that includes "unlimited calls" loses money on the customer running 5,000 calls a month. The vendor either caps usage (which is functionally per-minute), eats the loss (unsustainable), or raises the flat price to absorb the heavy tail (which prices out the average customer).
- Feature incentive misalignment. A flat-rate vendor has financial incentive to limit longer calls, deeper integrations, larger context windows. The product roadmap drifts toward the cheapest version of itself.
- Long-tail experimentation. Some buyers test heavily and then run experiments forever. A vendor on a flat rate is funding the experiment.
Pure flat-rate works for vendors with a single tight use case, strict usage caps, and pricing power over the buyer (enterprise contracts). It does not work for SMB self-serve at any reasonable price point.
The hybrid models that actually work
The pricing models that survive on both sides are hybrids. Four worth knowing:
1. Tiered flat with overage. Buyer picks a tier with a published minute or call cap. Overages run at a published per-minute rate, typically lower than spot per-minute. This is the dominant pattern in the AI voice category at SMB ($59 / $210 / $599-shaped tiers with overages at $0.08 to $0.15 per minute) and matches what most B2B SaaS uses for usage-based pricing in 2026.
This works because:
- Buyer gets predictable cost up to the tier ceiling.
- Vendor gets margin protection on the heavy tail.
- Overage pricing is published, not surprised.
- Tier upgrade math is explicit.
2. Outcome-based pricing. Buyer pays per qualified outcome (qualified lead, booked appointment, signed retainer). Vendor takes the volume risk and the qualification risk. Buyer takes only outcome cost.
This works in verticals where outcomes are countable and high-value (legal intake, life insurance, mortgage origination). It does not work where outcomes are diffuse (general customer support, broad triage).
The tradeoff: outcome pricing usually runs 2 to 5x the per-call equivalent of per-minute, because the vendor absorbs conversion-rate risk. Buyers who can run their own conversion math should sometimes prefer per-minute or tiered-flat. Buyers who cannot should prefer outcome pricing.
3. Annual commit with flat monthly. Buyer commits to 12 months of usage at a published flat rate. Vendor offers a discount versus monthly. Buyer locks in price predictability in exchange for longer commitment.
This is the enterprise pattern. It works at scale and breaks below roughly $1,000 a month, where contract overhead is too high relative to deal size.
4. Per-seat for internal use. Buyer pays per agent or per workspace, with usage included. Best for outbound sales (per-SDR seat) and for internal triage (per support-team seat).
Most serious vendors mix two or three of these. Pure single-model vendors are usually either at the technical-infrastructure end (per-minute, no SLAs) or the legacy-enterprise end (annual flat, sales call required).
How to evaluate a vendor's pricing model
The diligence pass for pricing takes 15 minutes. Five questions:
- What is my predictable monthly cost at typical usage? A vendor who answers "depends on volume" without giving a band is a vendor whose pricing will surprise you.
- What happens at my heavy month? Specifically, if I 3x my call volume next month, what is my bill, and what is the vendor's behavior (silently bill, throttle, contact me to upgrade)?
- What is the tier upgrade math? A vendor without published tier math is a vendor who plans to negotiate every pricing conversation.
- What is the cancellation policy? Month-to-month vs annual lock-in vs auto-renew. The auto-renew language matters most. Read it.
- Is outcome pricing available, and what does it require? Outcome pricing is the right answer for buyers in countable-outcome verticals. If the vendor does not offer it at all, they are not serving that buyer profile.
Vertical-specific pricing considerations
Different verticals reward different pricing models. Five worth naming:
- Real estate brokerage. Volatile call volume by season. Tiered flat with overages is the right fit. Per-minute punishes the slow months that subsidize the busy months.
- Legal intake. High value per qualified case, lower call volume. Outcome pricing works well (per qualified case, per signed retainer). Per-minute also works because volume is bounded.
- Insurance. Seasonal cycles (open enrollment, renewal). Tiered flat with annual commit makes sense for established carriers; per-minute for small agencies.
- Healthcare clinic intake. Steady call volume, white-label channel partners common. Flat-rate per clinic with bundled minutes is the right fit. Per-minute breaks the partner economics.
- Outbound sales (B2B). High call attempt count, low conversion rate. Per-seat or per-minute both work; outcome pricing usually does not because attribution is too diffuse.
The right pricing model for a vendor is the model that matches the verticals they serve. A vendor selling primarily to legal-intake firms should offer tiered flat plus outcome. A vendor selling to white-label partners should offer flat-rate per end-customer with bundled volume. A vendor selling to developers should stay per-minute. A vendor that offers all three at once is either very mature or unfocused. Usually unfocused.
The buyer-side decision framework
The clean buyer-side decision tree:
- Is your call volume predictable within plus or minus 25% month over month? If yes, flat-rate or tiered-flat is fine. If no, you need overage pricing or outcome pricing.
- Is your outcome countable and high-value? If yes, get an outcome-pricing quote and compare to per-call equivalent. If no, stick with tiered-flat.
- Are you reselling? If yes, you need flat-rate at the vendor layer or you cannot quote at the resale layer. Walk if the vendor only sells per-minute.
- Is annual commitment cheaper for you than month-to-month flexibility? Below $1k a month, usually no. Above $5k a month, usually yes.
Most buyers default to whatever pricing model the first vendor pitches. That default is wrong for at least half the buyer profiles. The pricing-model fit is more decisive than the per-unit price.
For other vendor-evaluation reading, see what AI voice agents fail at in the first 30 days, conflict checks at intake, and the after-hours math. To see how a tiered-flat plus overage model runs in practice, the pricing page lays it out.
