Why Contact Center AI Pilots Fail to Scale, and What Successful Teams Do Differently
Most contact center AI pilots stall for operational reasons, not technical ones. Here's why, and what the teams that reach production do differently.

Make CX Current News one of your go-to sources on Google
Contact center AI pilots tend to stall because of how they are scoped, integrated, tested and measured. The usual causes are a first use case that is too broad, integration and data gaps, agents designed from scripts instead of real conversations, testing that stops at launch and success measured by deflection alone. Teams that reach production start with a narrow use case, build on their own conversation data and keep evaluating the AI after launch.
The gap between pilot and production
MIT's Project NANDA reported in 2025 that only about 5% of integrated AI pilots were delivering millions in value, with the large majority showing no measurable effect on profit and loss. The authors traced the problem to brittle workflows, tools that couldn't learn from context and poor fit with day-to-day operations. Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, and analyst Anushree Verma added that "many use cases positioned as agentic today don't require agentic implementations."
Metrigy CEO Robin Gareiss said in a 2025 discussion with Cresta that her firm's research found about 53% of companies already showing a return from AI in customer experience. Customer service has clear success measures and high volumes of repeatable work, which gives AI projects a better chance than in many other functions. Even so, plenty of contact center projects still stall, and the causes tend to repeat from one company to the next.
Why contact center AI pilots stall
The first use case is too broad
Ambitious pilots that try to automate a large share of contact volume at once tend to run into every edge case simultaneously. In Gareiss' view, many projects fail because "too many companies try and boil the ocean." A pilot scoped to a handful of well-understood request types can be measured, fixed and expanded, which is much harder when the scope is all of customer service.
Integrations and data aren't ready
Order systems, billing platforms and CRMs that lack clean APIs leave an AI agent able to discuss a customer's problem without being able to fix it. In Cresta's 2026 CX Workforce Report, 81% of the 300 leaders surveyed named integration complexity as the biggest barrier to AI adoption, and only 7% said they could easily access their own conversation data. Aditya Challapally, lead author of the MIT study, said enterprise AI needs careful integration with existing systems and that without it, "pilots rarely scale."
The agent is designed from scripts instead of real conversations
Many pilots start from standard operating procedures and idealized call flows, which real customers rarely follow. Cresta's applied research team has written that designs built only from scripts and SOPs tend to harden into brittle flowcharts, little more than a fancier IVR, and recommends grounding an agent's scope in historical transcripts so it reflects what customers actually ask and how successful human agents resolve it.
Testing stops at launch
A pilot that passes its acceptance tests can still degrade once it meets real traffic, a new model version or an updated policy. Cresta has observed that most AI agent testing programs peak on launch day, after which requirements drift and test coverage falls behind the agent. Without ongoing regression testing and monitoring of production conversations, problems surface through customer complaints instead of before them.
Success is measured by deflection alone
When the business case rests on how many contacts the AI kept away from people, a pilot can hit its target while customers leave unhappy, and leadership loses confidence once the problems show up elsewhere. In the same CX Workflow survey, 89% of leaders reported gains in productivity and decision-making from AI, which a deflection-only business case would miss.
What teams that reach production do differently
Cresta customers that have moved AI from pilot to full deployment tend to follow a similar sequence, starting with work that has a clear payoff, building on their own conversation data and expanding in stages as results come in.
Brinks Home: building the foundation first
Brinks Home started with Cresta in 2023 by giving its human agents real-time AI guidance, then used what it learned from those conversations to deploy an AI agent across voice and chat. By the time automation went live, the company already understood what drove resolution and customer satisfaction in its own calls. Veronica Moturi, then Brinks Home's SVP of customer experience, said the result let the company automate "a wider range of conversations than we ever thought possible." Cresta reports that the multi-year program has contributed to a 30-point increase in Brinks Home's Net Promoter Score.
Propel Holdings: starting with a few high-value use cases
The fintech lender Propel Holdings needed to handle rising volume without growing headcount at the same pace. Cresta worked with Propel's operations team to pick a short list of starting points, including account management, payment inquiries and application support, where automation could show value quickly. Propel's AI agent for chat reached a 58% containment rate, and AI-generated summaries cut after-call work from three minutes to 90 seconds. The company is now extending the deployment to additional brands and teams. President and COO Gary Edelstein said Cresta was already "driving faster, smoother interactions" for Propel's teams and customers.
United Airlines: proving value inside the pilot window
United began with generative AI guidance for its contact center agents and set specific goals for the pilot. The airline reported surpassing those goals within the first 45 days, including a 15% reduction in average handle time. United has since built on that start, combining real-time agent guidance with conversation insights from Cresta to inform faster policy and customer experience decisions.
Xanterra Travel Collection: tying automation to revenue
Xanterra, which runs lodging in several U.S. national parks, reports a 74% containment rate and a $3.3 million revenue lift from its Cresta AI agent. Because the project was measured against revenue as well as cost savings, its value was easy to demonstrate to leadership.
A practical path from pilot to production
Choose the first use case from your own conversation data. Look for high-volume request types with clear resolution paths and low risk if something goes wrong. Cresta's applied research team uses topic analysis and reconstructed conversation flows from historical transcripts to size the opportunity and risk for each candidate before any agent is built.
Define success in business terms before launch. Agree on verified resolution, customer satisfaction, revenue or agent time saved as the measures, and set the targets the pilot has to hit to earn a wider rollout.
Fix integrations and data access early. Identify every system the agent needs to read from or write to, and confirm those connections work before the pilot starts.
Test against real customer behavior and keep testing after launch. Simulated customers built from historical conversations catch edge cases that scripted tests miss. Cresta's engineering team also reruns regression tests after every model or prompt change, which in one case caught an agent that had stopped reading a required compliance disclaimer.
Plan the human side of the rollout. Decide how handoffs will work, who will supervise AI agent conversations and how frontline agents will be trained for the harder calls they'll take on. Cresta's Agent Operations Center, for example, lets supervisors monitor live AI agent conversations and step in when needed.
Expand in stages. Add new use cases, channels or brands one at a time, using results from the previous stage to make the case for the next.
Cresta's guide to scaling AI in CX, based on lessons from leaders at Aptive Pest Control, Aqua Finance and TailorCare, reaches the same conclusion and argues that small, visible wins build the momentum larger programs need.
Frequently asked questions
Why do AI pilots fail in contact centers?
The most common reasons are a first use case that is too broad, systems the AI can't integrate with, agent designs based on scripts instead of real customer conversations, testing that ends at launch and business cases built only on deflection.
What percentage of AI pilots make it to production?
Estimates vary by study and industry. MIT's Project NANDA found that only about 5% of integrated enterprise AI pilots were producing millions in value in 2025, and Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027. Contact center results tend to be stronger, with Metrigy reporting that about 53% of companies were already seeing a return from AI in customer experience.
How do you choose the first AI use case for a contact center?
Start with high-volume requests that have clear resolution paths, reliable system access and low risk if the AI makes a mistake. Historical conversation data is the best guide to which requests fit that description.
How long should a contact center AI pilot take?
It depends on scope, but a narrow pilot should show measurable results within weeks. United Airlines reported beating its pilot goals within 45 days.
How do you measure whether an AI pilot is ready to scale?
Verified resolution, customer satisfaction, accuracy, agent productivity and revenue impact give a more reliable picture of whether the pilot is working than containment does on its own.




