Glossary and Where to Start
Terms used across these notes, defined plainly, and the routes through the collection for the common situations.
Reference · Reference
Adverse impact — a process that is neutral on its face producing different pass rates by group. Measurable stage by stage, and most organisations have never measured theirs.
The principles in “Glossary and Where to Start” become easier to maintain when ownership and time spent are visible. Teams evaluating work-hour tracking software can use it to coordinate the administrative work around assessment, identify stages that repeatedly expand and plan capacity, while keeping the hiring decision tied to job-relevant evidence instead of raw activity totals.
For an independent benchmark, compare the process with U.S. OPM assessment resources; the useful test is whether the local method remains job-related, proportionate and explainable.
Calibration — assessors applying the same criteria the same way. Checked by having two people score the same response independently.
Criterion validity — whether assessment scores relate to later performance. The strongest evidence and the hardest to gather.
Critical incident — a specific story of somebody handling a situation notably well or badly. The raw material for structured interview questions.
Job analysis — a written description of what the role requires and what evidence would demonstrate it. The document everything else derives from.
Reliability — whether the assessment gives the same answer twice. An unreliable instrument cannot be valid.
Rubric — the scoring guide: criteria, a short scale, and a description of what each level contains.
Structured interview — the same questions, in the same order, scored against written criteria. Predicts substantially better than conversation.
Work sample — a task resembling the job. The best-supported method available.
Terms used loosely elsewhere
"Validity" is not a property of a test. It is a property of using that test for that purpose, which is why nobody can sell it to you.
"Potential" usually means an impression formed from familiarity. Define what you mean or stop using it.
"Culture fit" frequently means similarity to the people already there, and it is where a great deal of bias enters.
And any precise validity coefficient should be treated with caution: recent re-examinations found earlier figures were inflated, and the ordering between methods is more reliable than the numbers.
Where to start
Designing a process from scratch: the job analysis, then which method predicts what.
You have a process and it is too long: how long an assessment should take, and the audit.
Hires are not working out: what assessment cannot do, then the audit.
Worried about fairness: adverse impact, then experience proxies.
Assessors disagree: training assessors, then calibration.
Candidates complain: candidate communication, then take-home tasks.
If you read only three
Prediction and what the evidence supports, because it determines what to build.
Job analysis, because everything derives from it.
And common failures, because you probably have three of them.
A closing note
Nothing here is legal advice. Discrimination, data protection and automated decision rules differ substantially by jurisdiction and change.
No product names, because the useful distinctions are between methods rather than between suppliers.
And no precise validity figures, for the reason given above.
What the collection argues
Assessment predicts performance to the extent that it resembles the work.
Organisations mostly assess what is convenient to measure, and the gap between that and what determines success is where bad hiring lives.
Every additional stage reduces one kind of error and increases another, and almost nobody makes that trade-off deliberately.
And the things that make a process fair, defensible and predictive are the same things: a written analysis, structure, evidence, and the willingness to measure your own outcomes.
The point
What makes a process fair, defensible and predictive is the same thing: a written analysis, structure, evidence, and measuring your own outcomes..
Underlying all of this
Almost everything in this collection reduces to one discipline: write down what the job requires, assess that thing directly, record the evidence, and look at your own outcomes afterwards. None of it requires buying anything, and organisations that do those four things consistently outperform ones running longer processes built from instruments chosen before the requirements were known.
Also in this section
Start here
Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.