When AI Handles the Easy Calls, Traditional CX Metrics Don't Mean What They Used To
Antony Passemard, VP of Customer Strategy at Cresta, on the rubber-band nature of customer experience metrics and why good AI can make traditional KPIs look worse.

Make CX Current News one of your go-to sources on Google
You have to think about the whole picture. And the whole picture is, how much does it cost you to properly serve your customer overall, at the level of satisfaction and quality you have set for your company?
Nearly every customer experience team has a favorite metric it's endlessly trying to improve. The problem is that optimizing any single metric in isolation produces a distorted picture, and AI makes the distortion sharper rather than clearer. A leader watching one number in isolation may conclude that an initiative like AI has failed, when in fact it worked exactly as intended. The only honest way to measure the CX function is across the whole operation at once, and that requires abandoning the comfort of a single headline number.
Antony Passemard looks at metrics from a different vantage point. As VP of Customer Strategy at Cresta, he helps enterprises design AI strategies that hold up against real business outcomes. A former Google Cloud executive with a background spanning AWS and Salesforce Service Cloud, Passemard has spent years inside the measurement debates that determine whether a company's AI investment is deemed a success or a failure. He cautions that the metrics most teams trust are the ones most likely to mislead them, first and foremost being average handle time, or AHT.
"As soon as you introduce AI agents, if they're really good, all the easy calls are going away. Your AHT goes up, because the humans only get the most complex and tough calls," he explains. The counterintuitive result of KPIs seemingly worsening when outcomes are improving is why single-metric thinking falls apart the moment AI enters the picture.
Pull one lever, create tension somewhere else
Passemard's mental model for customer experience metrics is a rubber band: the measures are interconnected, so improving one by force tends to strain another. "You pull somewhere, there's tension somewhere else," he says. "You can't have all of them at 100%. At some point something's going to give."
This is why he pushes back when a customer arrives fixated on a single target. The most common version is a leader determined to cut average handle time, who plans to deploy agent-assist AI toward that one number. "I have customers say, 'This is very important to us. They're spending too much time on the phone,'" he shares. In these cases, Passemard's caution is that the tactic can backfire in a way the customer doesn't anticipate. AI clears out the short, and more straightforward calls more reliably, which leaves the human team handling a queue made up disproportionately of the long, complicated ones, and the average climbs even though every individual interaction is being handled well.
The lesson isn't that handle time is worthless, but that the meaning of any metric in isolation may be misleading. The question that actually matters, in Passemard's view, sits at a higher level. "You have to think about the whole picture. And the whole picture is, how much does it cost you to properly serve your customer overall, at the level of satisfaction and quality you have set for your company?"
Cost to serve over calls per hour
Reframing around total cost to serve for a given quality (CSAT/NPS) pulls a range of measures into a cohesive view that a single metric obscures. Efficiency is only one dimension. The others Passemard recommends analyzing are about whether the operation is actually producing value. "Is the top-line revenue growing? Are your product defects going down? Is your margin or bottom line improving? Is your upsell improving? Is your first-call resolution improving? What about first-contact resolution, which is not necessarily the same thing as first call resolution?"
That distinction between first-call and first-contact resolution is a small example of the larger trap. Two metrics that sound like the same thing measure different realities, and a team optimizing the wrong one can believe it's improving while the customer experience degrades. The only defense, Passemard says, is to look at the outcomes together and accept that some will move in unwelcome directions even when the overall result is positive.
The four types of conversations
Passemard's organizing framework sorts every conversation into four buckets, and the framework matters because each bucket moves the metrics in a different direction. Treating them as one undifferentiated pile is what makes metric-chasing so misleading.
The first bucket is the conversation that shouldn't happen at all. "Your website or mobile app doesn't offer information customers are looking for, so they call you," he illustrates. "If I have 2,000 calls about my opening hours, let's put that on the website, and then they don't call."
The second bucket is the conversation worth automating: cases with good data, good APIs and where customers often prefer a fast automated resolution to waiting for a human. "Customers many times don't want to talk to somebody, they just want something solved. Those are the calls you should automate," he notes.
The third is the conversation that should stay human, where emotion, stakes, or VIP status make a person essential. "If I get in a car accident, I'm stressed out and I'm calling my insurance. I don't want to talk to an AI agent. I want to talk to somebody who's going to say, 'Hey, is everybody all right? Don't worry, we're going to take care of you.' That's your customer relationship. That's your brand."
The fourth and final bucket is the conversation you should be having but aren't: after-hours calls, proactive outreach, appointment confirmations, and the interactions that never happen today because there's no capacity for them.
Together, the framework predicts exactly the metric chaos that ambushes single-number thinkers. "If you apply an AI strategy across all four buckets, the metrics will change in various directions. Some will get worse, like AHT. Some will get better, like first-call resolution. You have to look at the overall output, keeping in mind a target CSAT or NPS, not individually in silos," Passemard advises.
As AI reshapes the space, Passemard warns that this problem will only intensify, because the mix of calls humans handle keeps shifting underneath the metrics. A number that meant one thing last quarter can mean something different this quarter purely because the composition of the queue changed. Any team still anchored to a single KPI will keep misreading its own performance, congratulating itself when a number moves for the wrong reason or panicking when a number worsens because the AI is doing its job. The discipline he argues for of judging the operation as a whole is the only way to know whether the customer experience is actually improving.





