Scoring Rubrics That Interviewers Can Apply
The document that turns judgement into assessment. What makes one usable, and the common forms that are not.
Designing · Procedure
A rubric that assessors cannot apply consistently is decoration. Most are, and the failures are predictable.
Once the criteria described in “Scoring Rubrics That Interviewers Can Apply” are fixed, the practical challenge is running them consistently across assessors. A coordinator can use this supporting guide to plan review time, compare the effort spent at each stage and spot avoidable delays, without confusing administrative efficiency with the quality of the evidence collected from candidates.
For an independent benchmark, compare the process with CIPD selection methods guidance; the useful test is whether the local method remains job-related, proportionate and explainable.
What a usable rubric contains
Three to five criteria, no more.
A scale with a small number of points — three or four, not ten.
A written description of each level, in terms of what the answer contains.
And an example of each level where possible, which does more for consistency than any amount of description.
Describing the levels
Bad: "excellent, good, fair, poor". These are labels, not criteria.
Bad: "demonstrates strong communication skills". Nobody knows what that means in practice.
Good: "describes a specific situation, their own actions and the outcome, and identifies what they would change".
The test is whether two people reading it would agree about a borderline answer.
Scale length
Three or four points.
Longer scales produce false precision: nobody can reliably distinguish a six from a seven, and the extra points add noise rather than information.
An even number forces a choice rather than allowing everybody to sit in the middle, which is worth considering where assessors are conflict-averse.
Weighting
Avoid it unless you have a reason.
Weighted composites feel rigorous and usually reflect somebody's guess about relative importance.
Equal weighting across criteria drawn from the job analysis is defensible and simpler, and where one requirement genuinely dominates, make it a threshold rather than a weight.
Thresholds versus ranking
A threshold says: below this, the candidate cannot do the job.
A ranking says: this candidate scored higher than that one.
Thresholds are defensible on most criteria. Rankings on small differences are not, because the instrument's resolution does not support them.
Use thresholds to shortlist and reserve ranking for genuinely large differences.
The free-text box
Always include one, and require evidence rather than opinion.
"What did they say that supports this score" is the prompt that works.
These notes are what you will need if a decision is challenged, and they are what makes calibration possible.
Testing it before use
Write three sample answers — strong, borderline, weak — and have two people score them independently.
Disagreement on the borderline one is expected. Disagreement on the strong or weak one means the rubric is not working.
An hour, before the first candidate.
What to check
How many criteria and how many scale points do you have?
Is each level described in terms of what the answer contains?
Do you require a note justifying each score?
And has the rubric been tested on sample answers by two people?
The point
Three to five criteria, three or four scale points, each level described in terms of what the answer contains.
Longer scales produce false precision.
Underlying all of this
Almost everything in this collection reduces to one discipline: write down what the job requires, assess that thing directly, record the evidence, and look at your own outcomes afterwards. None of it requires buying anything, and organisations that do those four things consistently outperform ones running longer processes built from instruments chosen before the requirements were known.
The recurring pattern
The recurring failure across every section here is the same: measuring what is convenient rather than what matters, then never checking whether it predicted anything. The check is an afternoon of work once a year, and it is the step that separates a process that improves from one that merely persists.
Also in this section
Start here
Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.