Skip to main content

AI PHONE & LEAD RESPONSE · September 2026 · ~11 min read

What callers actually do when they realize it is AI

Most of them keep going, as long as the system is fast and gets them what they called for. They leave when it stalls, loops, or refuses to hand them to a person. Realizing it is AI is rarely the moment a call fails. It is the moment the caller starts keeping score.

Owners worry about the wrong half of this. The fear is that a caller hears a synthetic voice and hangs up in disgust. What actually happens is that the caller notices, decides to give it a few seconds, and then judges it entirely on whether it is useful.

That is a much better position than it sounds. It means the voice matters far less than the competence, and competence is the part you control.

01

Do callers hang up as soon as they hear an AI?

Some do. Not many, and the ones who do were often going to be difficult calls anyway.

What I see in transcripts is a pattern rather than a cliff. The caller pauses, sometimes says "is this a real person," and then continues if the answer arrives quickly. The call ends when the agent fails, not when it is identified.

There is a second group worth naming: people who are relieved. Someone calling at eleven at night with a broken water heater does not care what is answering. They care that something answered.

The relevant comparison is not AI versus a great receptionist. It is AI versus your voicemail, which is what the caller was going to get.

02

What does the survey data actually say, and how much should you believe it?

It says people prefer humans, and it says so loudly. It also has a known weakness you should apply before you act on it.

A YouGov survey in January 2025 found that 55% of Americans prefer a human at the drive-thru, and only 4% prefer an AI chatbot. That is a wide margin and it is not noise. If you are building a system whose whole pitch is that callers will not mind, that finding is the argument against you.

Now apply the discount. This is a stated preference, not an observed behaviour. People are being asked what they would rather have in the abstract, with no cost, no wait, and no alternative described. Preference surveys in this genre consistently overstate how much the preference drives real choices, because the question never includes the tradeoff the person actually faces.

The tradeoff is the whole thing. Nobody in that survey was asked whether they prefer a human at the drive-thru or an AI that answers in four seconds when the human line is nine cars deep. Nobody calling your business at 9pm is choosing between AI and your best employee. They are choosing between AI and voicemail.

Read the 55% as a design constraint, not as a verdict. It tells you the caller starts out preferring not to be talking to you. It does not tell you they will hang up.

03

What do the big public deployments actually show?

They show that the failures were design failures, and that the survivor is the one that admits what it does not handle.

McDonald's ran automated order taking with IBM in more than 100 US drive-thrus from 2021 and ended the partnership in June 2024, with the technology switched off by 26 July. What went viral was not misunderstanding. It was an accumulation of long-tail modification errors: bacon added to ice cream, hundreds of nuggets on one order, items piling on after the customer said stop.

Taco Bell reached roughly 500 locations and publicly slowed its deployment in 2025 after viral trolling, most famously a customer who ordered 18,000 cups of water, and after around two million AI orders concluded that humans still belong in the drive-thru. That was adversarial input, a different failure from McDonald's.

Wendy's is the one still expanding, past 500 company-operated and franchise locations by late 2025, and reporting that 86% of transactions complete without employee intervention. Those are unaudited company figures and should be labelled as such. The interesting number is the other one. 14% of orders are deliberately routed to a human. Wendy's is not claiming full automation, and that is precisely why it is still running.

So the honest conclusion from the public record is not that consumers rejected voice AI. It is that two deployments shipped without quantity ceilings, absurdity detection or a clean escalation path, and one shipped with all three. Your line will be probed the same way within days of launch, by somebody who thinks it is funny.

04

What actually makes them leave?

Four things, in order of how often I see them.

Latency. The pause before the agent responds. Humans take turns at a modal gap of roughly 200 milliseconds, measured across ten languages by Stivers and colleagues in a 2009 study in the Proceedings of the National Academy of Sciences. Retell publishes an end-to-end target of 800 milliseconds at the median and 1,200 at the ninetieth percentile. Cross-vendor benchmarking by Cekura puts real-world median turn latency at 1.5 to 2.5 seconds.

Here is what that means on a real call. A twelve turn booking conversation at a 2 second median holds 24 seconds of silence. The same conversation between two people holds about 2.4 seconds. More than twenty seconds of your caller's experience is waiting, and they feel every one of them as something being wrong. Run the number on your own agent: count the turns in one transcript, multiply by your measured median, and you have the dead air budget you are asking a stranger to sit through.

Loops. The agent asks the same question again because it did not catch the answer. Once is forgivable. Twice and the caller concludes this is going nowhere.

Refusing to escalate. A caller asks for a person, gets told the agent can help, asks again, gets told the same thing. That is the pattern that turns a mild annoyance into a story they tell other people.

Being handled. A caller who suspects AI, asks directly, and gets a deflection feels tricked. They usually finish the call politely and never book.

Every one of those is a design decision, not a limitation of the technology. They are also the same causes behind most voice agents failing in their first week, which is not a coincidence.

05

Should you tell them upfront?

Yes, and be precise about why, because the legal question and the operating question are different.

There is no California statute requiring a general business phone line to disclose that a caller is speaking with AI. Rules are tighter in health care. AB 3030, effective January 2025, requires an AI disclaimer and instructions for reaching a human on generative patient communications, and it explicitly covers verbal ones. AB 489, in force since January 2026, bars any persona implying a licensed professional. Dental and med spa practices should treat both as counsel questions. Outside those settings, disclosure is your choice.

Make the choice anyway. A caller who was told at the start and then had a good experience has no complaint available to them. A caller who worked it out at minute three has a grievance, and grievances go to review sites rather than to you.

The wording matters more than the decision. "You have reached our automated assistant, I can book you in or get you straight to a person" does two useful things in one sentence. It is honest, and it tells the caller there is an exit, which is the thing they actually wanted to know.

And when someone asks directly, answer directly. Every time, no hedging.

06

Does the reaction change by time of day or by caller?

Considerably, and it should change what your system does.

At nine in the morning, a caller has options. Your competitors are open and answering, and their patience is short. During business hours, the bar for an AI is high, and the honest question is whether a person could have taken that call instead.

At nine at night, the alternative is nobody. Tolerance rises sharply, expectations drop, and the system is competing against silence. That is why after hours is the right place to start, and what after hours calls are actually worth is usually the strongest part of the business case.

The queue changes the calculation too. Patient Prism scored 11,552,668 dental calls across 8,280 locations in calendar 2025 and found 31 of every 100 callers hung up before reaching an agent at all. Those callers did not reject a machine. They rejected a wait. Against that baseline, an agent that answers on the second ring is not competing with a warm human voice. It is competing with hold music and a hang up.

Emergency callers are their own category. They want speed and a human, in that order, and the correct behavior is to capture the number and escalate immediately without qualifying anything. An agent that runs a burst pipe caller through five intake questions is doing real damage.

Seasonal spikes change the picture too. When demand concentrates into a few weeks, your competitors are also overwhelmed, and answering at all becomes the differentiator. Getting the site and the phone ready before a seasonal spike is worth more than anything you do during it.

07

What does this mean for how you build the thing?

Optimize for speed and exits, not for realism.

Do not chase a voice that fools people. It is expensive, it fails eventually, and the payoff when it works is smaller than the penalty when it does not. The YouGov number tells you the deception has a downside and no measurable upside.

Spend the effort on response latency, on interruption handling, and on a transfer that works the first time somebody asks. Those three determine the experience.

Give the caller a visible way out in the first fifteen seconds. Knowing an escape exists makes people less likely to use it. Wendy's version of this is a deliberate 14% human routing rate, and it is the reason its deployment outlasted two larger ones.

Build the ceilings before launch. Quantity limits, absurdity detection, rate limits, escalation after two failed turns, and a written policy for recorded abuse. McDonald's failed on the long tail and Taco Bell failed on the adversarial tail, and both were design gaps rather than model quality gaps.

And keep the phone in proportion. It is one of several ways people reach you, and the leaks between channels are usually bigger than anything inside a single call. Routing leads from four channels without dropping any is the wider problem. On mobile, a click to call button that is still missing costs more businesses more money than voice quality ever will.

08

What to do this week

Pull ten transcripts and find every point where a caller asked if this was a person or a machine. Read what happened next in each one.

Count the turns in one booking transcript and multiply by your measured median latency. That is your dead air budget, and it is usually the first thing to fix.

Rewrite the greeting to disclose and to offer the exit, in one sentence.

Try to break your own agent. Order 18,000 of something. Ask for a price nine times. Change your mind four times in one sentence. Log what it does.

Then call your own line and ask for a human immediately, twice. If it argues either time, change the rule today.

Be honest with yourself

When you do not need this

If your callers are almost entirely existing customers who know your staff by name, an AI on the main line will read as a downgrade no matter how well built. Use it after hours only, or not at all.

If your business runs on relationships and long conversations, the phone is your product and you should not automate it.

If your median latency is above two and a half seconds and your vendor cannot improve it, none of the rest of this matters. Fix the stack or do not launch.

And if your call volume is low enough that a person answers everything within two rings, nobody is going to realize anything, because there is nothing to realize. Leave it.

Sources

  • YouGov, January 2025 survey on drive-thru service preferences. Source of the 55% human preference and 4% AI chatbot preference. Consumer survey measuring stated preference, not observed behaviour. No stable public URL is recorded in our source library, so the citation is publisher and date only.
  • CNBC, "McDonald's to end AI drive-thru test with IBM," 17 June 2024. Business press. Source of the termination date and the 100-plus restaurant footprint.
  • Restaurant Dive, coverage of Taco Bell's drive-thru voice AI deployment. Trade press. Source of the 2025 slowdown, the roughly 500 location footprint and the trolling incidents.
  • Retell, published latency percentiles. Vendor documentation. Source of the 800 millisecond median and 1,200 millisecond ninetieth percentile targets.
  • Patient Prism, "The Dental Patient Access Report," 2 July 2026. 8,280 locations, 11,552,668 calls. Source of the 31 in 100 abandonment figure. Vendor research, dental only.
  • Stivers et al., "Universals and cultural variation in turn-taking in conversation," Proceedings of the National Academy of Sciences 106(26), 2009. Peer-reviewed source of the roughly 200 millisecond human turn transition gap.
  • Wendy's FreshAI performance figures, including the 86% autonomy rate, are company-reported and unaudited. Cekura cross-vendor benchmarks supply the 1.5 to 2.5 second real-world latency range and are vendor research.

Related reading

12

Questions about how your calls are landing?

Email me at eric@seod.com with one transcript or recording of an AI call from your line, with the caller's number removed. I will mark the exact moments where a caller would notice, where they would start losing patience, and where I would change the wording.

I read these myself. One call is usually enough to find the pattern, and the fix is usually two sentences in the script.

The rest of the AI phone and lead response library covers the pieces around the call itself.

Call Eric Email Eric