Skip to content
What You Are Measuring

All notes · Foundations

Prediction, and What the Evidence Supports

A century of research on selection has produced a few robust findings and a great many numbers that have not held up. What to rely on.

Foundations · Analysis

Selection research is unusually well studied and unusually badly summarised. The headline figures in circulation are less solid than their confidence implies.

The principles in “Prediction, and What the Evidence Supports” become easier to maintain when ownership and time spent are visible. Teams evaluating this time-management reference can use it to coordinate the administrative work around assessment, identify stages that repeatedly expand and plan capacity, while keeping the hiring decision tied to job-relevant evidence instead of raw activity totals.

For an independent benchmark, compare the process with U.S. OPM assessment resources; the useful test is whether the local method remains job-related, proportionate and explainable.

What holds up

Assessments that resemble the work predict performance better than assessments that do not. This is the most robust finding in the field and the one to build on.

Structure helps. The same interview, asked consistently and scored against criteria, predicts substantially better than the same conversation held freely.

Combining a small number of different methods beats any single one, with diminishing returns quickly.

And unstructured conversation predicts weakly while feeling highly informative to the person conducting it. That gap between confidence and accuracy is well documented.

What has not held up

The specific validity coefficients from older meta-analyses. A major re-examination in recent years found that corrections applied to the original figures had inflated them, sometimes considerably.

The result is not that selection methods do not work. It is that the ranking between them is more reliable than the numbers attached, and that everything predicts less strongly than the older figures suggested.

Treat any precise figure with caution, including ones quoted approvingly in vendor material, and rely on the relative ordering instead.

What the ordering looks like

Near the top: work samples, structured interviews, job knowledge tests.

In the middle: cognitive ability tests, situational judgement tests, structured reference checks.

Near the bottom: unstructured interviews, personality questionnaires used alone, years of experience, graphology and similar.

That ordering is stable across analyses even where the numbers move.

Why the weak methods persist

Unstructured interviews feel like the part where you learn the most.

Personality questionnaires produce a readable report, which is satisfying and looks like evidence.

Experience requirements are easy to apply and feel defensible.

None of those is about prediction, and recognising the reason is the first step to changing it.

What this means for your process

Include at least one assessment that resembles the work.

Structure whatever interviewing you do.

Use two or three methods, not six.

And stop quoting percentages to your stakeholders, because they will not survive scrutiny and the argument does not need them.

Reading a vendor claim

Ask what the criterion was. Predicting a supervisor's rating is not the same as predicting output.

Ask on whom, and whether the study was independent.

Ask whether the figure is corrected, and for what. That question alone tells you how carefully the claim was made.

And ask for adverse impact figures, which the fairness section covers and which vendors rarely volunteer.

What to check

Does your process include anything resembling the actual work?

Are your interviews structured or conversational?

Where did the numbers in your business case come from?

And could you defend your method ordering to somebody who had read the literature?

The point

Rely on the ordering between methods rather than on the numbers attached to them.

Recent re-examination found earlier validity figures were inflated by the corrections applied.

Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.