AI PHONE & LEAD RESPONSE · September 2026 · ~10 min read
Call volume, overage, and how AI phone pricing works
Almost every AI phone vendor charges a platform fee plus usage, and usage is billed in minutes. The number that decides your bill is average call duration, not call count. Ask for the overage rate, the rounding rule, and whether spam calls are billed, before you look at the headline price.
On this page
- 01How do vendors actually charge?
- 02What is actually inside a billed minute?
- 03What actually drives your minutes up?
- 04Why does the vendor's latency number not match the bill?
- 05What should you look for in the contract?
- 06How do you know if it is worth the money?
- 07What does a sane launch look like on the meter?
- 08What to do this week
- 09When you do not need this
- 10Sources
- 11Related reading
- 12Questions about a quote you were given?
Those three questions separate a predictable invoice from a surprising one. Most quotes lead with a monthly figure and an included minute bundle, and the bundle is sized on an assumption about your calls that nobody has checked.
I have watched this go wrong the same way repeatedly. The plan covers 500 minutes, the business is fine for three weeks, then a spam wave and one busy Monday burn the bundle and calls start failing or billing hard.
01How do vendors actually charge?
Four models, often combined.
Platform fee plus minutes. The most common. A monthly base, a minute bundle, and an overage rate per minute beyond it.
Per call. Simpler to forecast, and it punishes you for short calls. A fifteen second wrong number costs the same as a booked appointment.
Per resolution or per booking. Rare, and the most aligned with your interests. Worth asking for even if the answer is no, because the answer tells you how confident the vendor is.
Telephony passed through separately. Carrier minutes and number rental are sometimes billed on top of the AI usage. This line is small and it is frequently omitted from the quote.
Then there is buildout. Script writing, integrations, and testing are real work, and a vendor charging nothing for them is either doing it badly or recovering it in the usage rate.
Compare vendors on total cost per booked appointment, never on monthly price. A cheaper plan with longer calls and a worse script is more expensive on the only measure that matters.
02What is actually inside a billed minute?
Four separate services, each billing in a different unit, and this is the part that explains the invoice.
Vapi documents its own cost model plainly, and it is representative of the category. The transcriber bills per minute of two-channel audio. Text to speech bills per character. The language model bills per million tokens, and Vapi's own estimate assumes roughly five model requests per conversational turn and about 150 output tokens per minute of generated speech, blended against a prompt cache hit rate of around half. Telephony bills per minute on top.
Read that again, because there is a real operating consequence in it. A bloated system prompt is charged on every one of those five requests per turn, not once per call. Prompt length is a cost line, not just a latency one. Retell separately notes that prompts beyond roughly 8,000 tokens are noticeably slower, so the same bloat hits you twice.
The add-ons are small individually and they are not usually in the quote. Retell publishes aggressive background speech removal at half a cent per minute and its AI quality analysis at ten cents per minute analyzed after the first hundred minutes per workspace. On a few thousand minutes a month those are real numbers, and they are the ones that make an invoice differ from a proposal.
03What actually drives your minutes up?
Not the things owners expect.
Spam and robocalls. These are billed in most contracts, and they can be a meaningful share of inbound volume on a number that has been published for years. No published call taxonomy measures spam at all. Maple's breakdown of 1.2 million calls has reservations, hours, menu, orders, events, modifications and a 2% "other" bucket, which is the only place spam could sit and is implausibly small for it. Treat your spam share as unmeasured, count it yourself, and ask directly whether the vendor filters it and whether filtered calls are billed.
Latency. Every pause while the system thinks is billable silence. Retell's own published targets are 800 milliseconds end to end at the median and 1,200 at the ninetieth percentile. Cross-vendor benchmarking by Cekura puts real-world median turn latency at 1.5 to 2.5 seconds. On a two minute call with forty turns, an extra half second per turn is twenty seconds of paid dead air, and it is invisible in a demo where you are listening to the words rather than the gaps.
Loops. A call where the agent misunderstands and re-asks runs two or three times as long as it should. These are also your worst calls, so you are paying a premium for your failures.
Booking. A booked appointment takes longer than an hours question. Maple's average call duration across its dataset is one minute forty-five seconds, and a booking conversation runs well past that. Switching booking on changes your average duration immediately.
Rounding. Per second billing and per minute billing produce very different invoices when a lot of your calls are short. Ask which one you are on. If it is per minute with rounding up, every wrong number costs you a full minute.
04Why does the vendor's latency number not match the bill?
Because most published latency figures exclude the parts that produce the delay.
Vapi states explicitly that its displayed latency metric excludes endpointing and transport time, and that endpointing can account for a meaningful share of perceived delay. Endpointing is the system deciding the caller has stopped talking. It is a deliberate wait, it happens on every turn, and it is not in the number on the slide.
The default settings make this worse. Deepgram's endpointing parameter defaults to 10 milliseconds, while production deployments typically run 300 to 500. A system left on the default interrupts callers constantly, which produces loops, which produces minutes. A system tuned to 500 milliseconds adds half a second of deliberate silence to every turn, which also produces minutes. Both are billing decisions and neither appears in a pricing page.
So when a vendor advertises 500 milliseconds, that is not what the caller hears and it is not what your invoice reflects. Ask for measured end-to-end latency at the ninetieth and ninety-ninth percentiles on your own traffic, not the median, and not a lab figure. The median hides exactly the calls that hang up.
05What should you look for in the contract?
Six things, and you can ask all of them in one email.
The overage rate per minute, stated in dollars, not as "standard rates apply."
The rounding rule, per second or per minute.
Whether spam and abandoned calls are billed, and whether filtering costs extra.
What happens when you exceed the bundle. Does it bill through, throttle, or fail calls. A plan that silently fails calls at the cap is worse than no plan, because you are paying to be unreachable.
Contract length and what happens to your phone number, scripts, and call recordings if you leave. Number portability in writing.
And whether the price changes when your volume grows. Seasonal businesses get hurt here, and if your category has emergency search spikes that concentrate demand into a few days, your busiest week is exactly when you cannot afford calls to fail.
06How do you know if it is worth the money?
Worked example
Model it before you sign, using your own numbers. Here is the full walk-through.
Step one, size the traffic. Say 300 inbound calls a month, of which a human already handles 180 well. The system is for the other 120 calls.
Step two, estimate duration honestly. Maple's measured average is 1 minute 45 seconds across a large mixed dataset. Booking conversations run longer. Use 2.5 minutes as a working figure once booking is switched on. 120 × 2.5 = 300 minutes a month.
Step three, add the noise. Assume a tenth of your inbound is spam that gets answered and billed before anything filters it, at roughly 20 seconds each. That is another 12 calls and about 4 minutes. Small, until the wave hits.
Step four, price it against the bundle. A plan with a 500 minute bundle covers you at 304 minutes with headroom. A plan with a 250 minute bundle puts you 54 minutes into overage every month, and the overage rate is the number nobody quoted you. Ask what 54 minutes costs before you sign, not after.
Step five, run the value side. Of those 120 calls, apply your own close rate. ServiceTitan's platform data gives 24% for a shop with fewer than five technicians. 120 × 0.24 = 28.8 jobs. At a $400 average ticket that is $11,520 a month, and even after halving it for calls you would have recovered anyway, the coverage clears almost any subscription.
Step six, compare the whole cost. Platform fee, usage, telephony, add-ons, buildout amortized over twelve months. If that total is not comfortably smaller than the discounted value figure, do not proceed.
For low ticket businesses this arithmetic often fails, and the honest answer is a cheaper tool. An automatic text back on a missed call costs a fraction of a full voice agent and recovers a meaningful share of the same calls.
The same discipline applies to paid traffic, and the conversion math that decides whether ads make sense uses the same three inputs. If you cannot produce close rate and average job value, that is your first project, not this one.
What does a sane launch look like on the meter?
Start narrow and let the meter teach you.
Route after hours only for the first two weeks. Volume is low, risk is low, and you get a real average duration from your own callers rather than a vendor estimate.
Then add the overflow hours where you know you miss calls. Then, if the numbers hold, everything.
Watch average duration weekly during that ramp. If it climbs, read those transcripts. Rising duration is almost always loops rather than deeper conversations, and loops are a script fix rather than a billing problem.
Watch prompt length too. Every knowledge base addition is a permanent cost increase on every call, charged five times per turn. The discipline is to move reference material into a retrieval step rather than pasting it into the system prompt.
Do the duration testing before you go live. Testing an AI phone agent before it takes real calls gives you a rough duration figure for free, and it is a better forecast input than anything on a pricing page.
If you are still deciding what the system should be doing at all, what an AI receptionist actually does is the place to start before the pricing conversation.
08What to do this week
Get your inbound call count and, if your system reports it, average call duration for the last thirty days.
Email your shortlisted vendors the six contract questions in one message, plus two more: what does your published latency figure exclude, and what is your endpointing setting. Compare the replies rather than the pricing pages.
Build a one page model using the six steps above: expected calls, expected duration, spam allowance, bundle, overage, add-ons, buildout, total. Put your discounted recovered revenue estimate next to it.
If the two numbers are close, do not sign. Close is a loss once reality adds fifteen percent to the cost side.
Be honest with yourself
When you do not need this
If your call volume is genuinely small, the whole pricing conversation is academic. Any plan covers you, and the buildout cost is the only real number.
If you have not yet measured how many calls you miss, none of this modeling is real. Measure first.
And if your inbound is dominated by existing customers with account questions, an AI agent is an expensive way to annoy people who already pay you. Spend the money on a better callback habit.
Sources
- Vapi, "Understanding cost". Vendor documentation. Source of the per-stage billing units, the roughly five model requests per turn assumption, the 150 output tokens per minute figure and the statement that its displayed latency metric excludes endpointing and transport time.
- Retell, latency documentation. Vendor documentation. Source of the 800 millisecond median and 1,200 millisecond ninetieth percentile targets, the 8,000 token prompt threshold and the published add-on rates.
- Deepgram, endpointing documentation. Vendor documentation. Source of the 10 millisecond default, against production values of 300 to 500 milliseconds.
- Maple, "The state of restaurant phone communication". 1.2 million calls, 1,000+ US locations. Source of the 1 minute 45 second average duration and the 2% "other" category. Vendor platform data.
- ServiceTitan, call booking rate data, June 2022. Source of the 24% booking rate for shops with fewer than five technicians. Vendor platform data.
- Cekura cross-vendor voice agent benchmarks. Source of the 1.5 to 2.5 second real-world median turn latency range. Vendor research from a company selling voice agent testing.
Related reading
- AI booking into your existing calendar. Switching booking on is what moves your average duration, so read it before you size a bundle.
- Measuring your actual missed call rate before you fix it. Step one of the model above depends entirely on getting this number right.
- After-hours calls and what they are worth. The narrow launch described here is also the cheapest way to prove the business case.
- What to review monthly once an AI agent is answering. Average duration and prompt length belong on that monthly list, because both drift upward quietly.
Questions about a quote you were given?
Email me at eric@seod.com with the pricing page or proposal you received, plus your monthly call count. I will translate it into cost per answered call and cost per booked appointment at your volume, and flag which line items are missing from the quote.
I do this myself, no sequence attached. The missing line items are usually telephony, add-ons and buildout, and they are usually the reason the real invoice does not match the proposal.
More on the decisions around it sits in the AI phone and lead response library.