Situational Judgement Tests
Scenarios with response options, scored against expert consensus. Middling prediction, good candidate reception, and easy to build badly.
Methods · Analysis
A situational judgement test presents job-relevant scenarios and asks what the candidate would do. It sits between an interview and a work sample.
Once the criteria described in “Situational Judgement Tests” are fixed, the practical challenge is running them consistently across assessors. A coordinator can use the platform summary to plan review time, compare the effort spent at each stage and spot avoidable delays, without confusing administrative efficiency with the quality of the evidence collected from candidates.
For an independent benchmark, compare the process with CIPD selection methods guidance; the useful test is whether the local method remains job-related, proportionate and explainable.
What it is good for
Roles where judgement matters more than technical skill: service, care, supervision, client-facing work.
Candidates without directly relevant experience, since scenarios can be explained rather than requiring prior exposure.
Volume, because it scores automatically.
And candidate perception: people find them relevant, which reduces complaints and dropout.
How they are built
Scenarios drawn from the job analysis, preferably from real critical incidents.
Response options written to be plausible, including the common wrong answers.
Scored against consensus from experienced people in the role, gathered in advance.
That last step is what makes it an assessment rather than a quiz, and it is the step that gets skipped.
The common failures
Scenarios invented by somebody who has not done the job, which produce obvious right answers.
Options where one is clearly correct, which measures reading comprehension.
Scoring keys set by one person's opinion rather than by consensus.
And generic off-the-shelf scenarios, which have no relationship to your specific role and predict correspondingly less.
What they measure
Knowledge of what a good response looks like, which is not the same as doing it under pressure.
A candidate can identify the right answer and behave differently in the moment.
Which is the method's principal limitation and the reason it sits below work samples.
Should-do versus would-do
Asking what somebody should do measures knowledge.
Asking what they would do invites a more honest answer and a less consistent one.
Should-do formats predict slightly better and are easier to score. Either is defensible provided you are consistent.
Building your own
Feasible in a day for a specific role, and considerably better than a generic product.
Six to ten scenarios, four options each, consensus scoring from three or four experienced people.
Pilot it on current staff: if your best people do not score well, the key is wrong rather than they are.
That pilot is the step that separates a working instrument from a plausible one.
Accessibility
Long text-heavy scenarios disadvantage some candidates for reasons unrelated to the job.
Keep scenarios short, offer extra time, and check the reading level is no higher than the role requires.
What to check
Were your scenarios drawn from real situations in this role?
Was the scoring key set by consensus or by one person?
Have you piloted it on people already doing the job well?
And does it sit alongside a work sample, or is it carrying the decision alone?
The point
Pilot a situational judgement test on people already doing the job well.
If they do not score highly, the scoring key is wrong rather than they are.
Underlying all of this
Almost everything in this collection reduces to one discipline: write down what the job requires, assess that thing directly, record the evidence, and look at your own outcomes afterwards. None of it requires buying anything, and organisations that do those four things consistently outperform ones running longer processes built from instruments chosen before the requirements were known.
The recurring pattern
The recurring failure across every section here is the same: measuring what is convenient rather than what matters, then never checking whether it predicted anything. The check is an afternoon of work once a year, and it is the step that separates a process that improves from one that merely persists.
Also in this section
Start here
Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.