
Evaluating AI Phone Calls
Introduction
At Splash Aquatics, one of our core values is to surprise and delight our clients. We try to do this in little ways, such as remembering swimmer’s birthdays, checking in with every family at least once during a session, and being flexible in our policies for missed lessons.
However, as a small business, we face a significant operational challenge: we don’t have a dedicated Customer Support Representative (CSR). Instead, it falls to our administrative team to field phone calls between 8:00 AM and 8:00 PM, on top of our daily responsibilities. This sometimes results in periods where calls are missed, and not returned in a timely manner, which doesn’t align with our core value. To help us achieve a better answer rate, we’ve recently been piloting an AI phone agent tool called Retell AI. Before implementing the pilot, we analyzed the types of phone calls we were receiving and noticed the types of calls fit into 3 categories: informational, assistance, concerns.
Because retrieving and relaying structured information is something Generative AI excels at, we began piloting an AI phone agent to handle informational calls. By automating the routine informational inquiries, we ensured that the calls ringing through to our admins were the complex, emotionally nuanced ones that require genuine care.
While the operational efficiency of an AI agent is highly appealing for our admin team, the data suggests we needed to tread carefully. A 2025 ServiceNow study highlighted a stark reality about consumer preferences: 75% of respondents still want to talk to a human on the phone, and only 10% trust a voice agent to handle simple support tasks.
This tension between business efficiency and client hesitation is the focus of this evaluation. Can an AI successfully handle the nuances of our frontline support without alienating our clientele? To find out, I developed an evaluation heuristic grounded in AI ethics, affective computing, and sustainability, and put our Retell AI agent to the test.
Automating the Informational Flow
To start, I mapped out the flow of how I handle an informational call, trying to isolate patterns and codify best practices. An informational call is typically a prospective client looking to find out more about our specialized programs.
For this pilot I restricted the phone agent to informational to minimize risk, and to protect client data, the AI has no read access to our database. It only has write access to pass basic information (like an email or phone number) to a secure human-monitored inbox. Please check out the flow below:
The Evaluation Heuristic
To assess the AI's performance, I developed a rubric that evaluates a few different areas: Affective Resonance, Information Accuracy, Privacy, Transparency, and Sustainability.
As Taina Bucher notes, algorithms possess an affective dimension—they generate feelings and shape our emotional realities (2025). It is vital to evaluate the feel of the AI. Furthermore, as Lucy Suchman argues, we must avoid misplaced concreteness. LLMs do not truly understand our clients, they are “stochastic parrots” that are able to recite policies but not understand what they are, how they may affect different business functions, or impacts they may have on clients (2023). Therefore, the framework tests for factual adherence rather than perceived comprehension, alongside vital metrics for privacy (informed by UBC Guidelines) and environmental sustainability (2026).
Reflection
To test the heuristic, I roleplayed as a prospective client seeking information about our swim programs, specifically pushing the boundaries of the system by asking a series of follow-up questions about company policies, requesting unauthorized discounts, and attempting to trick the AI into making promises it couldn't deliver. I then began asking for specifics regarding signing up and rescheduling lessons on my account, tasks that the AI isn’t able to achieve. The results were largely reassuring on a functional level. The agent strictly adhered to our established policies and procedures, completely avoiding hallucinations or making unauthorized commitments. When confronted with queries it couldn't resolve, it defaulted to admitting its limitations and offered to transfer the call to a human admin. This aligned perfectly with the informational accuracy I was evaluating; the AI functioned well as a structured information retriever.
However, the AI's limitations became apparent in its affective responses and ability to empathize. During the test, I presented a scenario where my fictional child was sick quite often due to a medical condition and might miss several lessons. The AI immediately fell back to our policy of allowing only two make-up lessons per session. While technically accurate, this response directly contradicts our core value of surprising and delighting our clients. A human CSR might take a more flexible approach, perhaps offering an accommodation to ensure the family doesn't feel pressured to bring a sick swimmer to the pool. This highlighted a significant gap in both how we set up the policy within the AI phone agent’s knowledge base and how the AI prioritizes or understands the importance of client relationships.
Beyond the phone call itself, a few areas of the evaluation framework specifically related to the setup of the Retell agent. For example, to address the environmental impact of generative AI, I specifically selected the Gemini 3.0 Flash model as it is known for speed and efficiency but doesn’t consume as much computing power or energy as more sophisticated models like Claude Opus. To ensure transparency we hardcode the agent to introduce itself immediately as an AI customer service representative, explicitly stating its limitations and its ability to transfer the caller to a human at any time. Finally, to ensure strict privacy protections, we restricted the AI to write-only operations that collect the necessary information from the client and pass it along to a human to review and implement. We didn’t provide the AI agent with read access to our database to ensure that there was never a possibility of client data being accidentally leaked due to a misunderstanding from the LLM.
Going into this assignment, my initial assumptions were that I would mostly be concerned with the privacy of our database and the basic functionality of the agent. However, as I began mapping out the process, and thinking about how I would respond in specific situations I quickly realized there is a lot of nuance to the CSR role. As Bucher notes regarding the affective responses of algorithms, we have to consider how people's preconceived notions and feelings about AI shape the interaction before the bot even speaks (2025). Empathy plays a huge role in how I build trust and relationships with families on the phone and although AI can recite a cancellation policy and save our admin team time, automating the care required to support a neurodiverse community remains a human task.
References
Bucher, T. (2025). Beyond the hype: Reframing AI through algorithms and culture. Journal of Communication, 75(1), 81-84.
Carolino, B. (2025, May 14). Canadian consumers seek emotionally intelligent AI in Customer Service: Servicenow Study. Lexpert | Business of Law. https://www.lexpert.ca/news/technology-law/canadian-consumers-seek-emotionally-intelligent-ai-in-customer-service-servicenow-study/392492
Suchman, L. (2023). The uncontroversial ‘thingness’ of AI. Big Data & Society, 10(2) https://doi.org/10.1177/20539517231206794
Teaching & Learning Guidelines - Generative AI. The University of British Columbia. (2026, January). https://genai.ubc.ca/guidance/teaching-learning-guidelines/