Skip to main content
ARVATOBPO · LLC
All resources
Quality

Running a QA calibration that changes behaviour

Most calibration sessions produce agreement and no change. This is the agenda we run weekly — sample selection, blind scoring, variance review, and the single coaching commitment that leaves the room.

Playbook3 min read

Most calibration sessions end in agreement and change nothing. Six people listen to a call, discuss it for forty minutes, converge on a score, and go back to marking exactly as they did before. The session produced consensus on one call rather than alignment on a standard, and those are not the same thing.

Here is the agenda we run weekly. It takes sixty minutes and it is deliberately uncomfortable in two places.

Before the session

Sample selection is where calibration is usually lost. Pick the samples deliberately:

  • One call scored near the top, one near the bottom, and — the important one — one that scored in the middle, because the middle is where scoring disagreement actually lives.
  • At least one call that generated a customer complaint or a repeat contact. Outcomes the customer told you about are the best calibration material you will ever get.
  • Never let the person who scored the call originally choose it.

Send the recordings and the scorecard out at least a day ahead. Everyone scores independently, in writing, before the meeting. Scores are submitted to the facilitator and not shared until the room is together — otherwise the first person to speak sets the answer.

Minute 0 to 10: the spread, not the score

Open with the distribution, not the discussion. Put every participant's score for call one on screen at once.

The number you care about is the spread between the highest and lowest scorer, per criterion. A team that averages 87 with scores ranging from 71 to 98 is not calibrated, and the average conceals it completely. Track that spread week over week — it is the only real measure of whether calibration is working.

Minute 10 to 40: blind review of the widest gap

Take the criterion with the widest spread first. Not the call — the criterion.

Ask the highest and lowest scorers to each state, in one sentence, the evidence in the call that produced their score. Evidence means a moment, quotable, at a timestamp. "It didn't feel empathetic" is not evidence and should be sent back for rephrasing.

This is the uncomfortable part, and it is the whole point. Nine times out of ten the disagreement turns out to be about the scorecard's wording rather than about the agent — two analysts applying two reasonable readings of the same criterion. That is a finding. Write it down.

Minute 40 to 50: fix the scorecard, not the scorer

Every criterion that produced a wide spread gets one of three outcomes, decided in the room:

  • Reworded, with the new wording written on the spot and effective immediately.
  • Split, because it was measuring two things at once — accuracy and tone in one line is the classic offender.
  • Deleted, because nobody could articulate what behaviour it was driving.

A scorecard that survives a year of calibration untouched is not a stable scorecard; it is one nobody has argued about honestly.

Minute 50 to 60: one commitment

End with a single coaching commitment, named and dated. Not a list. One behaviour, one owner, one date by which it will show up in the next round of scores.

The reason it is one and not five is that five commitments produce nothing and one produces something. If the session cannot agree on which one matters most, the session did not reach a conclusion, and it is better to admit that than to record five items nobody will chase.

What to check a month later

  • Has the score spread narrowed? If not, the sessions are social rather than corrective.
  • Did last month's commitments show up in the scores? If a behaviour was coached and nothing moved, the problem is the process or the tooling, not the agent.
  • Are agents disputing scores less? Falling dispute volume after a rewording round is the clearest sign the criteria finally say what everyone thought they said.

Further reading

More on what running a QA calibration that changes behaviour means in practice.

Read on

Start the conversation

Want this applied to your numbers?

Bring what you have and we'll tell you honestly whether outsourcing your function makes financial sense — including when it doesn't.