Call centres
Quality scoring
Publish the rubric, calibrate the scorers, sample randomly, show the agent everything, and never rank anybody on it. For another implementation reference, see the Monitask overview.
Scoring works when the rubric is publicThe first condition
An agent who knows in advance what is being assessed, and how it is weighted, can meet the standard. One who receives a number afterwards learns only that they were judged.
Publishing the rubric costs nothing and it converts scoring from surveillance into a standard, which is the difference between a programme that improves calls and one that produces resentment and a stable average. Teams comparing the wider operating context can also consult UKG.
The scorers disagree more than you expect
Research on rating in general finds that a large share of the variation in scores comes from who is doing the scoring rather than from what is being scored. There is no reason call scoring escapes this.
Which makes calibration the most valuable and least popular part of a scoring programme: several scorers assess the same calls independently, the differences are examined, and the rubric is clarified where they diverge.
An organisation that has never calibrated does not know whether its scores mean anything, and cannot find out from the scores.
Sampling
Random within a stated frame, not selected. A sample chosen by somebody looking for something measures what they were looking for.
State the sample size, keep it consistent, and accept that a small sample of an agent's month is a noisy estimate. Reacting to a single bad score is reacting to noise, and it teaches agents that the programme is arbitrary.
What the agent seesDisclosure
The score, the rubric it was assessed against, the reasoning, the call itself, and a route to disagree that produces an answer from a person.
Agents see every score about them, with its reasoning and the call it came from.
Disagreements are recorded and answered. A scoring programme with no appeal is a rating exercise rather than a quality one.
Do not rank people on it
A published league table converts a diagnostic into a target. Agents optimise the rubric, scores rise, calls do not improve, and the measurement stops informing anybody.
Use scores for coaching individuals and for spotting patterns across the team. The moment they appear in a comparison, the entire programme has been spent.
Automated scoring, and its limits
Speech analytics can reliably establish that words were said, that a required disclosure was read, and roughly how the talking was distributed. Those are useful and they are mechanical.
What it does not establish is whether a call went well. Sentiment inferred from audio is a weak signal and it is frequently sold as a strong one. A number attached to how a customer felt, produced without asking them, is an estimate wearing the clothes of a measurement.
Use automation for the checkable parts and people for the judgement, and say which is which in the rubric.
Coaching
Scoring exists to make calls better. If nobody sits with an agent afterwards, the programme is a measurement exercise and the scores will not move.
Which is worth costing before starting: the scoring is the cheap part and the coaching time is the expenditure that makes it worth anything.
Where to start
Score twenty calls against a draft rubric with three people scoring independently, then compare. The disagreements will show you what the rubric does not say clearly, and fixing that before launch is the cheapest version of calibration available.