There’s a shift happening in CX and contact centers. AI agents were already in use in many of these workplaces, taking care of tasks like automating interactions and summarizing conversations.
But increasingly, AI is also being used to evaluate the people who are working in these centers by completing tasks like identifying coaching opportunities, scoring conversations and recommending when a manager should intervene in conversations.
In short, AI is now evaluating employees. But who is evaluating these AI evaluators?
In workplaces using these AI-enabled evaluation agents, employees may not be clear on how these assessments are happening and what they need to do to challenge them when they disagree with the evaluations. Workers not only need to know that these automated evaluations are happening in the first place, they also need to be aware of the measures in place to ensure those evaluations are fair, relevant and accurate.
Some companies are working to address this problem – for example, Microsoft now promotes a Quality Assurance Agent that can evaluate both AI and human interactions at scale. As both the use of these AI evaluators expands and firms introduce tools to help ensure those evaluations are appropriate and useful, quality assurance could become more comprehensive as more interactions are evaluated – as long as they are evaluated correctly.
More coverage isn’t necessarily better evaluation
Automated QA holds several potential benefits for CX and contact centers. Traditional QA approaches for these centers rely on sampling a relatively small number of overall interactions. Using AI, however, means a much wider sample of interactions can be evaluated, well beyond the quantities humans could reasonably tackle.
But it’s important to understand the distance between coverage and independently verified quality. Simply double-checking a random sample of AI-generated assessments may not be sufficient, says Michał Piszczek, chief technology officer at Archdesk. Piszczek instead advocates for risk-tiered, stratified validation, with evaluations sampled across categories like interaction type, shift, language and team. At the same time, assessments with a direct effect on things like employee pay, performance reviews, or discipline should get independent human review.
The amount of scrutiny an AI assessment is subject to should depend, at least in part, on what the organization plans to do based on the results of that assessment. “Tie the verification budget to the stakes of the decision, not to call volume,” Piszczek said.
Ultimately, a system can be both very consistent and consistently incorrect. An AI-enabled evaluator might produce QA scores that make sense according to multiple measures, even while the relationship between those scores and the actual performance of the CX or contact agents being measured is deteriorating.
How do you evaluate an evaluator?
Continuous testing of AI-generated assessments against independent human judgement is one way to address that problem of evaluating an AI evaluator. For example, Oura manually validated more than 3,000 interactions before deploying its Total Customer Experience (TCX) system to evaluate customer interactions, says John Moses, Oura’s VP of member experience. With the system in place, QA specialists independently audit more than 1,000 interactions each quarter and compare their assessments with TCX on measures like recall and precisions, Moses says.
Classifiers or prompts are revisited if performance falls below the company’s threshold and the evaluations are re-run. At the same time, disagreement between the evaluation methods isn’t necessarily seen as a failure, Moses says. “The important distinction is that AI provides the scale, but humans continue to define the standard,” he said.
Oura’s model highlights several concrete practices that can be put in place in other situations where AI-enabled evaluators are in use. Their TCX evaluates nearly every interaction, with human QA experts benchmarking the assessments made. Where there is disagreement between the two, a human employee investigates them to find out why it happened. Was there context missing? Was the prompt too broad? And when those investigations identify discrepancies between the human- and AI-led evaluations, the system can then be recalibrated accordingly.
These questions highlight why organizations shouldn’t move directly from introducing an AI evaluator to relying on its assessments. An incremental deployment approach is better, says Aler Rab, deputy CEO of Cloudzy. Systems should be tested internally, and their conclusions compared with existing human assessments. The role of the AI evaluator should be expanded only when its reliability has been established.
“Trust in AI should develop in much the same way trust develops between people: through repeated evidence of reliability, not by assumption,” Rab said.
Some things are harder to score than others
Some AI evaluation systems are being asked to measure qualities of CX and contact interactions that are more complex than straightforward questions of right or wrong answers.
For example, there are meaningful differences between questions like “Did the agent follow the prescribed process?” and “Was the agent empathetic?” As the interactions move down the spectrum from the former to the latter, assessment becomes increasingly contextual and subjective.
Evaluating the evaluator also means examining what it’s been asked to score and which criteria it’s been given, Rab says. To do this, managers should be involved before an AI system is deployed, rather than simply receiving its scores after the fact. “AI should be built around the organization’s understanding of good performance, not the other way around,” he says.
That understanding itself can shift once AI evaluators are part of a workplace, because evaluation systems can change the behavior they’re meant to measure. Once workers understand the scoring logic, they might start to optimize their performance for that rubric, Piszczek says. In this ways, rising QA scores alongside resolution metrics that hold steady could be a sign of a measurement problem, not of improved performance.
Over about a year spent developing TCX, Oura considered more than a dozen potential signals to track before choosing just four: issue identification, frustration, resolution and post-solution sentiment. The company then looks for evidence that those four measures correspond to meaningful CX problems, Moses says.
For example, Oura discovered its agents were routinely asking customers to repeat information they already provided. After a human review of the issue, they found it stemmed at least in part from workflows, scripts and handoffs between chatbots and humans – not necessarily with the representatives themselves, Moses says. This highlights how an AI system can correctly identify a bad customer experience but may not reveal who or what caused the issue.
The Oura example illustrates why a poor customer interaction doesn’t necessarily mean an employee performed poorly. Throughout the scoring process for AI evaluators, the key question moves from asking if the model applied the rubric correctly to ensuring the criteria was appropriate to evaluate in the first place.
How workers can help evaluate the evaluators
“The worker’s right to challenge a score isn’t just fairness. It’s your only freees to know they are being evaluated by AI-enabled systems. They need to be able to see the assessment, understand the basis for it, have a means to challenge it if desired and be assured a human review – one that can change the outcome of the assessment, if warranted – can be accessed
Those challenges can also become part of the validation process itself. Both dispute and overturn rates should be tracked across teams and segments, Piszczek advises. If a cluster of successful challenges of AI assessments, around a specific criterion of evaluation, is found, that might mean the issue lies with the criterion itself.
It’s also important to understand that an absence of challenges shouldn’t be automatically interpreted as evidence of the accuracy of the evaluations. If workers aren’t aware they can contest their AI assessments, don’t understand how to do so, or are convinced their challenges will have no impact, the absence of challenges is no indication of the reliability of the AI-enabled evaluations’ results.
“An employee challenging an AI assessment is not a failure of the system. It is another data point for improving the system,” Rab said. At Cloudzy, original CX interactions, their operational context, employee reporting and AI-generated assessments are all retained. Using that information, managers can investigate a disputed assessment using multiple sources instead of treating the AI assessment as definitive, he says.
Ultimately, detecting a problem using AI-enabled assessments is not the same as determining the responsibility for the problem. AI might detect a poor experience, but humans need to remain in the loop along the way: to investigate the cause, to determine responsibility and to determine and apply appropriate interventions.
“At critical points, AI can expand the evidence available to a manager without becoming the manager,” Rab said. AI-enabled evaluation can be a valuable part of the coaching and assessment process but shouldn’t overtake it entirely.
Evaluation needs its own feedback loop
Often, with AI-enabled workplace interventions, it’s important to ask if a human appears somewhere in the process. But for AI evaluations, the question should go a step further to investigate if an organization has created an effective feedback loop around the evaluator itself.
That loop contains several key steps. It begins with AI assessment of an interaction before moving to independent validation of that assessment, which should be judged for accuracy and fairness by a manager. Workers should have the ability, if needed, to challenge those assessments. And finally, the AI-based assessment model should be recalibrated or corrected when any step of this process identifies an issue.
When faults are found with an AI evaluator, it’s important to consider the impact of the decisions its already made. To do this, Piszczek suggests an approach used by engineers dealing with a defective software release: determine the blast radius of the error using timestamps and model versions, allowing an organization to identify and revisit consequential decisions made by the AI evaluator before it was course corrected. “Forcing workers to absorb the cost of a defective QA model is bad engineering and bad governance,” he said.
As useful as it can be for some workplaces, automated QA creates a paradox. Using AI, organizations can evaluate a much larger quantity of their employees’ work than ever possible before. And that very ability means the humans evaluating the AI evaluators are more important, not less. When a human evaluator makes an incorrect call on an interaction, that affects one assessment. But when an AI evaluator gets it wrong on a call related to empathy, quality, or efficacy, that mistake can quickly scale across thousands of interactions.
“Scale does not distinguish between a good metric and a bad one – it amplifies both,” Rab said.
