AI PHONE & LEAD RESPONSE · September 2026 · ~11 min read
What to review monthly once an AI agent is answering
Six numbers and ten transcripts, once a month, in about forty minutes. The numbers: calls answered, average duration, booking rate, escalation rate, failed transfers, and cost per booked appointment. The transcripts: five escalated calls and five short ones. Everything you need to fix will be in there.
On this page
- 01What numbers should you look at every month?
- 02What does a real three month trend look like?
- 03Which calls should you actually listen to?
- 04Which metrics are being gamed, and what should replace them?
- 05What do you actually change as a result?
- 06What needs a quarterly check instead of a monthly one?
- 07What else belongs in the same review?
- 08What to do this week
- 09When you do not need this
- 10Sources
- 11Related reading
- 12Questions about your monthly review?
An AI phone agent is not a purchase, it is a system, and systems drift. The script that fit your business in March does not fit it in September, because your prices changed, your hours changed, and callers started asking about something you did not offer six months ago.
The businesses that get real value from this category are not the ones that bought the best product. They are the ones that reread it every month.
01What numbers should you look at every month?
Six, and they take ten minutes to pull.
Calls answered. The volume the system is actually handling. Compare against the same month last year if you have it, because seasonality will otherwise read as performance.
Average call duration. Your cost driver, since usage is billed in minutes. A rising average is almost never deeper conversation. It is loops. For a rough sense of scale, Maple's dataset of 1.2 million calls across more than 1,000 US locations puts the average local business call at 1 minute 45 seconds.
Booking rate. Appointments booked divided by calls answered. This is the health number. If answered calls rise and bookings do not, the system is answering and losing, which is worse than voicemail because voicemail at least left you a number.
Escalation rate. Both directions are a problem. Very high means the script does not cover your real calls. Near zero means the agent is answering things it should hand off, which is invisible until a customer complains.
Failed transfers. Calls where a human was requested and nobody picked up. This is the metric nobody measures and it is the one that costs you customers, because a stranded caller had already asked for help.
Cost per booked appointment. Total cost including platform, minutes, and telephony, divided by appointments booked. This is the only number worth comparing against alternatives.
Write them in the same place every month. A spreadsheet with six columns beats any dashboard, because a dashboard shows you this month and a spreadsheet shows you the trend.
02What does a real three month trend look like?
Here is the arithmetic with invented numbers so you can see the shape. Rerun every line with yours.
Month one. 412 calls answered, average duration 1:52, 118 booked, 34 escalations, 5 failed transfers, total cost $640.
Booking rate is 118 of 412, or 29%derived. Escalation rate is 34 of 412, or 8%derived. Transfer failure is 5 of 34, or 15%derived. Cost per booked appointment is $640 divided by 118, or $5.42.
Month two. 455 answered, 121 booked. Booking rate 27%derived.
Month three. 470 answered, 119 booked. Booking rate 25%derived.
Now read it. Answered calls rose 14%derived across the quarter. Bookings did not move. Cost per booked appointment climbed because you are paying for more minutes to produce the same number of jobs.
That is the single most common failure pattern in this category and it does not look like a failure on any dashboard. The volume chart is going up and to the right. The business is not.
Three things to check when you see that shape, in order. Whether the extra calls are spam, which is easy and free to test. Whether the extra volume is a new marketing channel bringing worse leads, which is a marketing question rather than a phone question. And whether a script change in month one quietly broke a booking path, which is why you note the date of every change.
The 15% transfer failure rate deserves its own note. It is not a rounding error. One in seven people who asked for a human got nothing, and every one of them had already decided your agent could not help them.
03Which calls should you actually listen to?
Ten, chosen deliberately, not sampled at random.
Five escalated calls. Every escalation is a place your system reached its edge, which makes them the highest signal recordings you have. Look for the same question type appearing repeatedly, because that is a script gap you can close in an afternoon.
Five very short calls. Under twenty seconds usually means a hangup, and hangups tell you about latency, greeting, and voice. Nobody reviews these because they seem uninteresting. They are the ones showing you why people leave.
Add one compliance spot check every month. Pick a call where a caller pushed on a boundary and confirm the agent held it. In a dental practice, that means no statement that a procedure was covered and no out of pocket figure. In a restaurant, no answer to an allergen question. The full never say list should be checked against real calls, not against the script, because the script is what you wrote and the call is what happened.
Keep a running log of questions the agent could not answer. That log is next month's script update, and it is the most valuable document the system produces.
04Which metrics are being gamed, and what should replace them?
Containment rate, and the fix is to stop reading it alone.
Containment is the share of calls resolved without a human. It is the easiest metric in the category to improve dishonestly, because making escalation harder raises it. An agent that buries the transfer option will post a beautiful containment number and a worsening business. Always read it next to escalation rate, transfer success, and whatever customer satisfaction signal you have.
Word error rate is the other one to distrust. Deepgram's own explainer works through a phone call with a 25% word error rate, which it describes as about average for off the shelf recognition, where the transcript was entirely usable except that it turned "declined" into "designed" and the call was silently misrouted. Ask instead for the entity error rate: how often the name, the number, the address and the appointment type come out wrong. Those are the errors that cost money, and at least one platform reports them separately as mistranscribed entities.
Latency reporting has the same problem. A median is comfortable and useless. Retell publishes end to end targets of 800 milliseconds at the median and 2,500 at the ninety ninth percentile. The ninety ninth percentile is where the hangups live. Ask for p95 and p99 monthly, not the average.
And the strongest ground truth signal is not native to any platform I have seen. Correlate your call log against your CRM and count how many contained calls produced a callback from the same number within 24 hours. If a meaningful share of your successfully handled callers ring back the next day, the containment number is fiction. That report has to be built. It is worth building once.
05What do you actually change as a result?
Three types of change, in this order.
Script gaps. Anything in the unknowns log that came up more than twice gets an approved answer added. Approved means signed by whoever holds the license or liability, same as the original buildout.
Facts that went stale. Hours, pricing ranges, service area, staff names, which services you currently offer. This is the most common cause of a good system giving wrong answers, and it is entirely self inflicted. Your intake documents are the source of truth, so the information the agent needed before it went live needs the same monthly pass.
There is a cost angle to this that owners miss. Vapi's own documentation explains that a voice pipeline bills the language model on roughly five requests per conversational turn, so every answer you add to the prompt is charged on every turn of every call, not once. A knowledge base that grows all year without pruning gets slower and more expensive at the same time. Cut two answers for every three you add.
Routing rules. Who takes transfers, at what hours, on what number. This breaks whenever staffing changes and nobody thinks to update it.
Change one thing at a time where you can, and note the date. If booking rate moves and you changed four things, you have learned nothing about which one did it.
06What needs a quarterly check instead of a monthly one?
Four things that move slowly and matter a lot when they move.
Compliance posture. Recording notice wording and placement, retention periods, and any disclosure requirement in your industry. California requires the consent of all parties before recording begins, and for health care adjacent practices AB 3030 and AB 489 both bear on what the agent may say about itself. The rules in this area are moving. A quarterly review with your own counsel is cheaper than finding out late, and nothing in this article substitutes for it.
Two items belong on that list specifically. If your system also texts, confirm that a spoken opt out on a call writes into the same suppression list as a texted STOP, because federal revocation rules treat any reasonable expression of the request as valid and give you a limited window to honor it. And if any part of your setup places outbound calls in an AI voice rather than answering inbound ones, flag that separately, because the FCC has held that AI generated voices are artificial under the TCPA and require the prior express consent of the called party.
Contract terms. Your bundle against your actual minutes, your overage rate, and whether your volume has grown into a different tier.
Integration health. Calendar writes, CRM records, and notifications. These break quietly during software updates, and the failure looks like a slow month rather than a broken integration.
Whether the system is still the right shape. If your volume, hours, or service mix has changed substantially, revisit what the agent should be doing at all rather than continuing to tune a design that fit an older business.
07What else belongs in the same review?
The phone is one part of a conversion path, and the monthly hour is a good time to look at the rest of it.
If more of your inbound moved to the web, the question is whether people are reaching a useful next step. Getting an estimate request instead of just a phone call changes what the phone has to carry, and often improves both.
And when a number moves and you cannot explain it, watch people rather than charts. The first thirty session recordings will usually tell you what a month of metrics could not, because they show you the moment somebody gave up.
08What to do this week
Build the six column spreadsheet and fill in this month, even if the data is incomplete.
Pull five escalated transcripts and five calls under twenty seconds. Read all ten in one sitting.
Ask your vendor for p95 and p99 latency and for an entity error count rather than a word error rate.
Start the unknowns log. One line per question the agent could not answer, with the date.
Put a recurring forty minute appointment in your calendar for the same day each month. If it is not scheduled, it will not happen, and an unreviewed agent is a slowly worsening one.
Be honest with yourself
When you do not need this
If no AI answers your phone, the transcript half does not apply. The six numbers still work for a human front desk, and most businesses do not track any of them.
If your volume is very low, monthly is too frequent. Do it quarterly and spend the saved time on demand.
If your cost per booked appointment is climbing quarter over quarter and the script is not the cause, the honest conclusion may be that the system is not earning its place. Cancelling is a legitimate outcome of a review.
And if you are going to look at the numbers without changing anything, do not bother. A review that never produces a script edit is a reporting habit, not an operating one, and it will make you feel informed while the system quietly drifts.
Sources
- Maple, "The state of restaurant phone communication". 1.2 million calls across 1,000+ US locations, December 2023 to November 2025. Source of the 1 minute 45 second average duration. Vendor platform data.
- Vapi, "Understanding cost". Vendor documentation. Source of the per stage billing units and the roughly five model requests per conversational turn, which is why prompt length is a recurring cost.
- Retell, published latency percentiles. Vendor documentation. Source of the 800 millisecond median and 2,500 millisecond ninety ninth percentile targets, and the reference implementation for reporting critical entity errors separately from raw transcription error.
- Deepgram, "What is Word Error Rate?". Source of the 25% average error rate for off the shelf recognition and the declined versus designed example. Vendor documentation from a speech recognition company.
- California Penal Code section 632.7 on all party consent for recorded calls; AB 3030, Health and Safety Code section 1339.75; AB 489, Business and Professions Code sections 4999.8 to 4999.9. State statutes, primary sources.
- FCC Declaratory Ruling 24-17, 8 February 2024, on AI generated voices under the TCPA, and FCC Report and Order 24-24, revocation rules effective 11 April 2025, codified at 47 CFR 64.1200(a)(10) to (12). Federal agency orders, primary sources.
Related reading
- Why most AI voice agents fail in the first week. Month one is a different review from month twelve, and this is what to look for first.
- How good implementations handle escalation to a human. Read it the first month your failed transfer count is not zero.
- How call volume and overage pricing work. Where a rising average duration turns into a bill you did not plan for.
- Measuring your actual missed call rate before you fix it. The baseline your monthly numbers should be compared against, and the reason to have captured it.
Questions about your monthly review?
Email me at eric@seod.com with your current review process, or tell me plainly that you do not have one. I will send back the exact six column tracker I would use for your business, with the thresholds I would treat as warning signs at your volume.
I put these together myself. Send a month of numbers with it and I will tell you which one I would look at first, and whether the trend says keep it or cancel it.
More on everything around the phone sits in the AI phone and lead response library.