Real-time AI coaching is turning customer interactions into a continuous stream of performance data. Systems can analyze wording, tone, sentiment and adherence to scripts as a conversation unfolds, then prompt an employee or alert a supervisor. That speed can make assistance more useful, but it can also erase the distance between coaching and constant evaluation.
In the meantime, demand is growing, with organizations ranking AI-enabled coaching platforms among their top three AI application priorities for the next 18 months at 24%.
That compares with 56% for broad productivity tools such as ChatGPT, Gemini and Copilot, according to a recent IDC survey.
However, close to half (48%) cited a lack of security and governance protocols as their top concern about AI-enabled work models.
Intent, consent define the line
Dr. Amy Loomis, IDC group vice president for workplace solutions, says the technology itself does not determine whether a system is supportive or intrusive.
“The line is intent. Employee coaching and workplace surveillance are two sides of the same technology,” she said. “Coaching built into the workflow, prompting someone at the moment they need it, gets adopted because it solves a problem an employee has.”
Justin Beals, CEO of Strike Graph, an AI-native governance, risk and compliance platform, places equal emphasis on whether employees understand and can examine the evaluation.
“The line is consent and purpose, not technology,” he said. “Coaching that the employee knows about and can see the same data on that their manager sees is assistance.”
He points out scoring that runs continuously, feeds a permanent personnel record and was never disclosed as evaluative is surveillance wearing a coaching costume.
Transparency therefore must cover more than the presence of AI. Employees need to know what is measured, how scores are produced, who sees them and whether they can affect pay, promotion or discipline.
Beals notes workers should see their own scores at roughly the same time as supervisors, rather than discover them later through a performance process.
Loomis says it’s important to start by disclosing to employees that AI is in the loop.
“Ideally, companies should also be able to honestly say that humans are, not in the loop, but in the lead,” she said.
Empathy scores are probabilistic, not facts
The hardest evaluations are also among the most subjective. Tone, sentiment and empathy can depend on culture, relationship history, the customer’s behavior and events outside the system’s field of view.
An employee who briefly steps away from a call, for example, could appear disengaged without the model knowing why.
“[Current systems] are not reliable enough to be treated as ground truth and that’s the part getting lost,” Beals said. “Sentiment and empathy inference are probabilistic judgments about human behavior, not measurements.”
He cautions treating a confidence score like a performance fact is the same mistake seen in AI-driven hiring tools: the model looks precise, so people stop asking how often it’s wrong.
AI bias and hallucination were cited by a third of respondents as one of their largest security concerns around AI-enabled work models in IDC’s survey.
Loomis cautions behavioral scores can conceal mistakes more effectively than other AI outputs.
“A hallucination in a document task looks wrong and gets caught,” she said. “A hallucination in an empathy score looks like a number. Unfortunately, almost nothing is built to catch that.”
Cultural context compounds the problem, with Loomis noting a model trained on explicit, low-context communication may struggle to interpret high-context exchanges.
Subtle behavior such as a microaggression can depend on social position and relationship history rather than a single word or vocal cue, making human review and a process for challenging scores essential rather than optional.
Decide who owns the data before deployment
Organizations must also define access and retention before turning the system on. Loomis notes Europe’s data protection regime, which treats performance-related personal data with heightened care, offers a reasonable baseline in other regions.
Employees who freely opt into coaching without consequences for declining should be able to use their own data as a development signal. HR may still learn from aggregated and anonymized results about where a team needs support without exposing individual workers.
Beals warns against allowing access to expand merely because the information is available.
“Access should be limited to the people directly responsible for coaching that employee, not stored indefinitely in a system HR, legal and eventually litigation can all reach years later,” he said.
From his perspective, the retention question is really a risk question.
“Every day that inferred emotional data sits in a database is another day it can be subpoenaed, breached or used for a purpose it was never collected for,” he said.
Loomis says retention should follow the original purpose: Data collected to inform one evaluation period has done its job when that period closes.
“Holding it longer erodes the trust the coaching program was built to earn,” she said.
Governance must reach performance review
Policies are meaningful only if they govern how scores are used. Loomis says organizations should treat governance as an operating discipline.
Senior leaders must explain the purpose of coaching tools and business and HR teams reviewing annually how data connects to performance evaluations, safety rankings and other enterprise metrics.
The majority of respondents in IDC’s May 2026 Future of Work survey reported either a fully operational AI center of excellence or governance work underway; however, Loomis cautions formal structure is not enough.
“There’s a big difference between having a framework and having transparency employees actually experience,” she said. “A framework that nobody explains isn’t going to protect anyone very well.”
Beals says organizations need a documented purpose limitation that prevents coaching data from quietly migrating into disciplinary decisions, a human review before any score produces a consequence and an audit trail covering the model, not only the employee.
“If you can’t audit why the system scored someone’s empathy low, you have no business letting that score touch their performance review,” Beals said.
