Eight Questions Before Putting AI on Your Phones
Over the past two years, we have built conversational AI for service businesses. Along the way, we have learned that the gap between an impressive demo and a genuinely useful product can be enormous.

It is now easy to connect a language model to a phone line. The difficult part is making it work reliably in a real business: when the calendar is messy, several customers call at once, jobs are urgent, and nobody has time to clean up mistakes afterwards.
If you are considering a voice agent — whether you are buying one or building your own — these are the questions worth asking.
1. Does it fit the way your business actually runs?
Answering a phone call is only the beginning. The real question is whether the agent can turn that conversation into useful work.
For most service businesses, a booking is rarely as simple as choosing the next empty calendar slot. There are fixed appointments, flexible arrival windows, urgent jobs, multi-day projects, travel time, and the inevitable last-minute change. A useful system needs to understand those realities rather than forcing every job into the same template.
It should also help you make better scheduling decisions. A free slot may technically be available, but it could create an unnecessary drive, leave an awkward gap in the day, or push your team into overtime. The best scheduling tools consider the broader day, not just whether a time is open.
It is also worth considering what happens after the call or the booking. Is the tool integrated with your quoting and job-management tools, or are you copying and pasting information across platforms? For us that part is not an afterthought: the job goes straight into the job-management tool you already run, with the customer’s details and call notes attached, live, during the call. Quoting works the same way — the assistant collects the detail during the call and leaves you a draft ready to send, or you can use the Chime Quotes iOS app to build a quote between jobs, by dictation or by talking to Chimey.
The same standard applies to quoting. Can the system use your actual pricing? Can it catch a misheard email address or an incorrectly spelled name? If a quote already exists, can it update it instead of creating duplicates? And can you use it from the road rather than waiting until you are back at a desk?
A good test is simple: after a week, has your admin reduced, or have you just acquired another system to monitor?
2. Does it work directly with the tools you already use?
A voice agent is only as useful as its connection to the systems that run your business.
Many products create their own calendar or job list. That may be quick to set up, but it often leaves you managing two versions of the truth. You have to remember which system has the current information, and staff can easily miss changes.
A better approach is to connect directly to your existing job-management system. When a customer asks for next Tuesday, the agent should check live availability in the calendar your team already uses. When the booking is confirmed, it should create the job there, including the customer details and call notes. That is how our ServiceM8 integration works.
Look beyond the integration logos on a vendor’s website. Ask whether the connection can both read and write data. Does it check availability during the call? Does it create or update the customer record? Or does it simply leave a note for someone to deal with later?
If you are building internally, plan for integrations to take more effort than the phone agent itself. They are where much of the real complexity lives.
3. How do they measure whether the agent is improving?
Conversation quality is difficult to judge casually. Reading a few transcripts and adjusting a prompt might feel productive, but it is not a dependable way to improve a system.
You can fix one part of a call and accidentally make another part worse. Perhaps the greeting improves, but the agent stops asking an important qualifying question. Without structured evaluation, those regressions are easy to miss.
A mature voice product should score calls against clear criteria and test changes against realistic simulated conversations before they reach customers. This is essentially regression testing for phone calls.
Ask vendors how they know the product performs better today than it did last month. If the answer is mainly that someone occasionally reads transcripts, that is a warning sign.
4. Can it balance speed, intelligence, and reliability?
Every voice agent is balancing three competing priorities: how quickly it responds, how well it reasons, and how consistently it follows instructions.
On a phone call, even a short silence can feel uncomfortable. Callers may interrupt, assume the line has dropped, or hang up. But faster models are not always the best at handling complex requests, and highly capable models can sometimes be too slow for natural conversation.
The experience depends on the full pipeline: hearing the customer correctly, recognising when they have finished speaking, deciding what to do, and responding quickly enough to keep the conversation flowing. There is more on how we approach that on our technology page.
There is another important distinction: a model can be capable without being consistently obedient. Even strong models can ignore a rule when a conversation becomes unusual or complicated. That is why prompts should not be the only safeguard. Important business rules need checks in the surrounding system as well.
5. Is the team still improving the underlying technology?
The voice-AI stack changes quickly. Models, speech recognition, synthetic voices, and infrastructure all improve at a rapid pace.
A product can fall behind if its technology choices are difficult to replace. If changing a provider means rebuilding core parts of the system, improvements become slow and painful.
Ask what the team has changed in the past six months, and why. A company doing serious product work should be able to point to specific examples: testing newer models, comparing speech-to-text providers on the same calls, or maintaining backup providers in case one service has an outage.
The best technology stack is rarely permanent. What matters is whether the team can adapt without disrupting customers.
6. What happens when every phone rings at once?
People can only answer one call at a time. Most businesses are used to that limitation, but it becomes particularly painful during busy periods.
A storm can trigger a rush of calls for roofers, plumbers, and electricians. A successful ad campaign can create a sudden spike in enquiries. Those are often the most valuable calls, yet they are also the calls most likely to be missed — which is exactly the problem Chime Overflow is built for.
A voice agent should be designed to handle concurrent calls without degrading the experience. That requires more than a good demo. It requires systems that can scale, manage queues, and behave sensibly if a third-party provider slows down.
Ask the straightforward question: what happens when ten, fifty, or a hundred people call at the same time? The distance between handling one polished demo call and handling a real peak period is substantial.
7. Does it eliminate work, or merely move it somewhere else?
The useful metric is not the number of calls answered. It is the amount of administration removed.
If a staff member still has to copy a caller’s details into the job system, create the customer record, write the call notes, and send a follow-up, then the AI has only solved a small part of the problem.
A well-designed system should carry information through the whole workflow: capture and process the conversation, update the customer record, create the booking, attach relevant notes, and trigger any necessary follow-up. In some cases, it should be able to prepare a quote while the customer is still on the call.
Try tracing a single enquiry from start to finish. Count every time a person has to handle information the customer has already provided. Each extra step is an opportunity for delay, error, or missed work.
8. Does it feel like it understands your business?
The strongest voice agents do not sound like a spoken online form. They sound like someone familiar with the business.
That means knowing the services you provide, the areas you cover, your hours, your pricing approach, and how you handle urgent work. It also means avoiding repetitive questions when the system already has the information it needs.
Getting there should not require weeks of filling in setup forms. A good product should do much of the initial work using the information already available about the business, then let you review and correct it. The aim is to start with a useful draft, not a blank page.
Over time, the agent should become more useful as it learns the details that matter. Customers notice when they are talking to something that understands the business. They also notice when it is simply following a generic script.
The common thread
All of these questions point to the same thing: the difference between having an AI conversation and achieving a useful business outcome.
A voice agent should not just answer calls. It should help turn calls into booked jobs, accurate customer records, timely quotes, and less admin for the people running the business.
That is the standard we build Chime Labs around. The technology is only valuable if it makes the day run better for the people using it.
Chime Labs is an AI-powered front-office and operations platform for service businesses — Chime Reception, Chime Overflow, Chime Coach and Chime Quotes.
See what this looks like on your own phones
Hear how Chime handles a real enquiry, books the job into the system you already run, and leaves the notes behind.