Calibration Between Assessors
Two people scoring the same thing differently is the commonest defect in assessment. How to measure it and narrow it.
Running it · Procedure
Calibration means assessors applying the same criteria the same way. Without it you are not running one assessment, you are running as many as you have assessors.
For a recurring programme built around “Calibration Between Assessors”, improvement depends on seeing how the work is actually carried out between review cycles. the full explanation can help teams organise assessor time and compare planned effort with delivery, but the resulting records should support calibration conversations rather than become an automatic judgement about an individual.
For an independent benchmark, compare the process with U.S. OPM assessment resources; the useful test is whether the local method remains job-related, proportionate and explainable.
The exercise
Take three to five real responses: a strong one, a weak one, and one that divided opinion.
Everybody scores independently, without discussion.
Then compare and talk through the differences.
An hour, and it finds more than any amount of rubric revision.
What you will find
Wider disagreement than anybody expects, particularly on the borderline case.
Different interpretations of the same criterion.
One or two assessors consistently more lenient or severe than the rest.
And criteria that turn out to mean different things to different people, which is a rubric problem rather than an assessor problem.
Fixing what you find
Where the rubric is ambiguous, rewrite that level with a concrete example.
Where an assessor is systematically lenient, tell them — privately, with the data.
Where disagreement is about what the job requires, go back to the job analysis, because that disagreement is upstream of the assessment entirely.
Frequency
Before any substantial round.
Quarterly for teams hiring continuously.
And after any change to the rubric or the role.
Assessors drift, and the drift is not detectable from inside.
Live calibration during a round
Where two assessors score the same candidate, compare before discussing.
A large gap is a signal to look again at the evidence rather than to split the difference.
And record both scores rather than only the agreed one, which gives you calibration data for free.
Panel discussions
The discussion should come after independent scoring, always.
Otherwise the first or most senior voice anchors everybody, and the panel produces one opinion with four signatures.
Chairing it properly: each person states their score and one piece of evidence before anybody argues.
Who calibrates whom
Everybody with everybody, over time.
A pair who always assess together calibrate to each other and drift from the rest, which produces a consistent sub-process nobody intended.
Rotate pairings where volume allows.
What to check
Have your assessors ever scored the same response independently?
How wide was the disagreement?
Do panels score before or after discussing?
And are both scores recorded, or only the final one?
The point
Have assessors score the same response independently and compare.
Disagreement is wider than anybody expects and it is the finding.
Underlying all of this
Almost everything in this collection reduces to one discipline: write down what the job requires, assess that thing directly, record the evidence, and look at your own outcomes afterwards. None of it requires buying anything, and organisations that do those four things consistently outperform ones running longer processes built from instruments chosen before the requirements were known.
The recurring pattern
The recurring failure across every section here is the same: measuring what is convenient rather than what matters, then never checking whether it predicted anything. The check is an afternoon of work once a year, and it is the step that separates a process that improves from one that merely persists.
Also in this section
Start here
Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.