Sampling vs scoring every support call: what changes for QA
Sampling scores a few calls per agent and assumes the rest look similar. Scoring every call drops that assumption: policy breaches, repeated complaints and uneven agents show up when they happen, and reviewers spend their time judging the calls that need a person instead of hunting for them.
Published · Updated
How QA sampling works today
In most support teams, quality assurance runs on a sample. A reviewer pulls a handful of calls per agent each week or each month, listens to them against a scorecard, and records a score for items such as the greeting, identity verification, empathy, accuracy of the answer and the closing. Those few scores become the agent's quality number for the period.
Sampling exists for a good reason: listening takes time, and a reviewer can only hear so many calls in a day. A small, random sample is a reasonable way to estimate how a large team is doing when listening is the only way to know.
The trouble is what the sample is then used for. A handful of scores is treated as a verdict on one agent, a trend on the team, and evidence for a policy decision, all at once. It was never large enough to carry that weight.
Where sampling breaks down
The gaps in sampling are not about effort. Reviewers can work hard and still miss what matters, because the calls that matter are rare and scattered.
- Rare but serious events slip through. A missed identity check or a promise the policy does not allow may happen on a few calls a month, and a small sample can easily contain none of them.
- Scores swing on luck. An agent whose sample happens to include two angry callers looks worse than a colleague who drew easy calls, and the difference says nothing about skill.
- Trends arrive late. When a new product issue drives a wave of complaints, the sample shows it weeks after the phones did.
- Disputes are hard to settle. An agent who disagrees with a score is arguing about a call nobody else has heard, with notes as the only evidence.
- Reviewers spend most of their time finding calls, not judging them. The hard part of QA becomes the search, not the assessment.
Measured and judged criteria are different
Before scoring every call, split your scorecard into two kinds of item. Some criteria can be measured: whether the agent talked over the caller, how much of the call the agent spent talking, how long the agent's longest uninterrupted stretch ran. These have a number, and the number is either right or wrong.
Other criteria need judgment: whether the agent explained the refund policy correctly, whether they acknowledged the caller's frustration, whether they offered a resolution the caller accepted. These are opinions about what was said, and they are only as good as the evidence behind them.
The two kinds deserve different treatment. A measured item should show the number and how it was computed. A judged item should show the words it rests on, so a reviewer or an agent can read the quote and agree or disagree with the judgment in seconds. A score that shows neither is a score you have to take on trust.
What changes for the QA team
When every call is scored, the reviewer's job moves from sampling to checking. Instead of choosing which calls to hear, reviewers start from the calls that were flagged: a policy item that failed, a low score, a caller who mentioned cancelling. They confirm or correct the judgment, and they spend their listening time where it changes an outcome.
Agent scores also become more stable. A score built on every call an agent took in a week is not at the mercy of which two calls were drawn. That makes it fairer to use in a review conversation, and easier for the agent to accept.
Calibration gets easier as well. When two reviewers disagree about a judged item, they can both read the quote it rests on and settle the question on the wording of the call, not on memory. Over time that sharpens the scorecard itself, because vague criteria are the ones reviewers keep disagreeing about.
Coaching improves too. A team lead who can see that an agent skips the verification step mainly on calls transferred from another queue can coach that specific moment, with the calls that show it, instead of repeating the whole verification script in a general refresher.
And the team learns about its customers, not only its agents. When every call is read, the reasons people call, the complaints that keep coming back and the promises agents feel pushed to make become visible across the whole floor. That is often the most useful thing QA can hand to the product and operations teams.
Building a scorecard that works on every call
A scorecard written for human reviewers often assumes the reviewer will fill gaps with common sense. A scorecard used on every call needs to say exactly what it means.
- Write each criterion so two reviewers reading the same call would give the same answer.
- Keep measured items and judged items separate, and weight them on purpose.
- Give each department its own scorecard; a billing queue and a technical support queue are not graded on the same things.
- Mark the few items that are policy, not preference, so a breach is treated differently from a weak greeting.
- Review the criteria reviewers disagree on most, and rewrite them until they stop disagreeing.
Where Callens fits
Callens reads every sales and support call, in Arabic and English, and cites the second of the recording behind every claim. Scorecards use each department's own criteria and weights; judged criteria show the quote they rest on, and measured criteria show the number.
The measured criteria are talk ratio, open questions, interruptions and the longest monologue, drawn from 22 conversation measures computed from word-level timings rather than guessed by a language model. Every model-produced line is checked against a real segment of the transcript before it is stored, and risks are typed, including churn and compliance, and graded low, medium or high.
For teams that want QA to reach other systems, signed webhooks fire on a policy breach or a low score, as well as on completion and high risk. Numbers roll up by company, department or person, so a QA lead can see the whole floor and still open the single call behind any score.
Key takeaways
- Sampling was a sensible answer to limited listening time, but a few scores cannot carry a verdict on an agent and a trend on the team at once.
- Rare, serious events such as a missed verification are exactly what small samples tend to miss.
- Measured criteria should show their number and judged criteria should show the quote they rest on.
- Scoring every call turns reviewers from searchers into checkers who spend their time where a person's judgment changes the outcome.
Questions
What is QA sampling in a call center?
QA sampling means a reviewer scores a small number of each agent's calls against a scorecard and treats those scores as representative of the agent's work for the period.
Why is sampling calls not enough for quality assurance?
Because the calls that matter most, such as policy breaches or escalations, are rare, and a small sample can contain none of them. Scores also swing on which calls happen to be drawn.
Does scoring every call replace human QA reviewers?
No. It changes their work. Reviewers stop searching for calls and start checking the flagged ones, confirming or correcting judgments and improving the scorecard.
Can Callens score support calls with our own scorecard?
Yes. Callens scorecards use each department's own criteria and weights. Judged criteria show the quote they rest on, and measured criteria show the number.
Go deeper
Read next
- Why review every sales call instead of a sample?Why sales managers who hear a few calls a week miss the patterns that decide deals, and what changes when every call is read.
- Why do call summaries need timestamps?A call summary without timestamps asks you to trust it. Why every line of a summary should point to the second of the recording it rests on.
See Callens read your own calls.
Tell us about your calls. We can show you Callens reading a week of them, on your own recordings, before you decide anything.
Talk to us