Contact Center AI
Every conversation, not a sample

Sriram Chakravarthy, CTO & Cofounder
Here is a number that should bother every contact center leader who sees it: 2%.
That is the percentage of customer interactions most contact centers review for quality. In some organizations, it's 5%. In a few well-resourced ones, it might reach 8%. The rest go unheard.
The industry has accepted this for so long that it no longer registers as a problem. It is treated as a constraint of the operating model, like weather. QA analysts can only listen to so many calls. Supervisors have other responsibilities. The math doesn't work. So you sample, you extrapolate, and you hope that the 2% you heard is representative of the 98% you didn't.
It isn't. It never was. And everyone in the contact center knows it.
The sampling illusion
Random sampling works in contexts where each sample is statistically identical to every other. It does not work in the contact center, because customer interactions are not identical. They vary by agent, by time of day, by call type, by customer segment, by product, by mood, by season. A 2% sample of a population this variable is not a window into the whole; it is a keyhole.
Consider what a QA analyst actually does with that 2%. They listen to a call. They score it against an evaluation rubric. They note where the agent deviated from the script, missed a compliance requirement, or failed to follow process. They enter scores into a spreadsheet or a QA tool. Then they move to the next call.
This process takes 15 to 30 minutes per interaction. An analyst reviewing calls full-time might get through 15 to 20 per day. In a contact center handling 5,000 calls daily, that is four-tenths of one percent.
The calls that get reviewed are chosen randomly or, more commonly, cherry-picked: a supervisor flags a complaint, a customer escalation triggers a review, an agent is already on a performance plan. The calls that nobody flags go unheard. The agent who is quietly excellent gets no recognition. The agent who is consistently mediocre but never terrible gets no coaching. The compliance violation that doesn't generate a complaint goes undetected.
This is not quality assurance. It is quality guessing.
What the 98% contains
The calls that never get reviewed contain everything the contact center needs to know and currently doesn't.
They contain the agent who found a better way to explain a confusing policy, a technique that resolves calls 40% faster than the scripted approach. Nobody knows, because nobody heard the call.
They contain the three-week window where a specific product generated twice the normal complaint volume, a signal that something changed in manufacturing or shipping or documentation. Nobody noticed, because the QA sample didn't happen to include enough of those calls.
They contain the pattern where customers who call between 6 and 8 PM get worse outcomes than daytime callers, not because the evening agents are less skilled, but because the knowledge base they rely on doesn't include information about a product line that generates disproportionate evening volume. Nobody connected those dots, because the data lived in unreviewed calls.
They contain compliance risks. Not the dramatic kind that triggers an immediate escalation, but the subtle kind: an agent who consistently skips a required disclosure, a process that has drifted from the approved script over months, a workaround that technically violates regulation but has become so routine that nobody thinks to flag it. These risks are invisible at 2% sampling. They are obvious at 100%.
The 98% is not noise. It is signal. It is the majority of the business, operating in the dark.
Let's talk about "bias"
When a supervisor selects calls to review, they bring judgment to the selection. That judgment is sometimes informed and fair. It is also sometimes influenced by which agents the supervisor already has opinions about, which shifts get more scrutiny, and which call types are easier to evaluate.
The result is that QA scores often reflect the biases of the reviewer as much as the performance of the agent. Two supervisors scoring the same call will frequently disagree. The same supervisor scoring similar calls on different days will sometimes score them differently. Agents learn this quickly. They learn that their QA score depends as much on who reviews the call as on how they handled it.
This erodes trust in the entire process. Agents view QA as arbitrary rather than developmental. Supervisors view it as a compliance obligation rather than a coaching tool. The feedback loop that QA is supposed to create, the loop between performance and improvement, breaks down.
Eliminating the human from the evaluation isn't about efficiency. It is about consistency. When every call is evaluated against the same criteria, by the same system, with the same standards, the score reflects the interaction. Not the reviewer.
From sampling to sensing
The technology to evaluate every interaction now exists. Not in theory; in production.
AI can score 100% of customer interactions, across voice and chat, against the same evaluation criteria that QA analysts use today. It can detect compliance violations, script adherence, sentiment shifts, resolution effectiveness, and coaching opportunities. It can do this in real time, not weeks later. And it can do it consistently, without the variance that comes from human reviewers having good days and bad days.
This changes QA from a sampling exercise into a sensing layer.
At 2%, QA answers the question: "Did this specific call go well?" At 100%, QA answers different questions entirely: "Which agents are improving? Which are declining? Which call types generate the most compliance risk? Which coaching interventions actually work? Where are our processes failing systematically rather than individually?"
These are the questions that drive operational improvement. They are unanswerable at 2%. They are obvious at 100%.
The feedback gap closes
The most immediate impact of evaluating every call is on the agents themselves.
In the traditional model, an agent handles 80 calls in a day. One of those calls might get reviewed. The feedback arrives two weeks later in a one-on-one meeting. By then, the agent barely remembers the interaction. The feedback is accurate but useless, because the moment where it could have changed behavior has long passed.
When every call is evaluated, the feedback loop compresses from weeks to hours. An agent finishes a call and can see how it was scored before the next one rings. They can see patterns in their own performance across the day, the week, the month. They can identify their own strengths and gaps without waiting for a supervisor to notice.
This is what agents have always wanted. In the original conversations that inspired this work, the most consistent request from agents was simple: "Give me feedback on every call, not just the ones you randomly pick." The constraint was never willingness; it was capacity. That constraint no longer applies.
The shift in what QA means
For decades, quality assurance in the contact center has meant oversight: checking whether agents followed the rules, catching mistakes, documenting failures. It has been, in most organizations, a compliance function more than a performance function.
When you can evaluate every interaction, QA becomes something different. It becomes the system that tells you how your entire operation is performing, in real time, across every agent, every channel, every call type. It becomes the early warning system for process failures, compliance drift, and emerging customer issues. It becomes the coaching engine that makes every agent better, not just the ones who happened to get reviewed.
The 2% model was never designed to do any of this. It was designed for a world where listening was expensive and scale was impossible. That world is gone.
The question is no longer "how do we review more calls?" It is "what becomes possible when we hear all of them?"
The answer is that quality assurance stops being an audit and starts being an operating system. One that runs on every conversation, not a sample.

