Skip to content
What You Are Measuring

All notes · Foundations

Reliability and Validity, Without the Statistics

Two ideas that decide whether an assessment is worth running, explained in terms you can apply without a psychometrics background.

Foundations · Explainer

These words appear in every vendor brochure and are rarely explained. Both are simple and both are checkable in your own process.

The principles in “Reliability and Validity, Without the Statistics” become easier to maintain when ownership and time spent are visible. Teams evaluating this supporting resource can use it to coordinate the administrative work around assessment, identify stages that repeatedly expand and plan capacity, while keeping the hiring decision tied to job-relevant evidence instead of raw activity totals.

For an independent benchmark, compare the process with U.S. OPM assessment resources; the useful test is whether the local method remains job-related, proportionate and explainable.

Reliability: does it give the same answer twice

If two assessors score the same candidate differently, the assessment is unreliable.

If the same candidate would score differently on a different day, for reasons unrelated to the job, it is unreliable.

An unreliable assessment cannot be valid, because an instrument that gives random answers cannot give correct ones.

This is the cheaper problem to fix and the one most organisations have.

How to check yours

Have two assessors score the same five candidates independently, then compare.

Substantial disagreement means your criteria are too vague to apply, not that one assessor is wrong.

Repeat after training and see whether the gap narrows. Its own note covers calibration.

Validity: does it measure what you think

An assessment can be perfectly reliable and measure the wrong thing.

A typing test reliably measures typing. Whether typing predicts success in the role is a separate question, and it is the important one.

Validity is not a property of the test. It is a property of using that test for that purpose, which is why a vendor cannot sell you validity — only evidence that it worked somewhere else.

The kinds you will hear about

Content validity: the assessment covers the actual work. A work sample has this by construction, which is why it performs well.

Criterion validity: scores correlate with later performance. The strongest evidence and the hardest to gather, because it requires following hires over time.

Face validity: it looks relevant to candidates. Not evidence of anything, and it matters anyway — candidates who find an assessment irrelevant disengage and complain.

What you can actually check locally

Whether your assessors agree.

Whether your assessment scores relate to anything about performance six months later, which requires recording scores and revisiting them and is the single most valuable thing most organisations never do.

And whether candidates who scored badly but were hired anyway turned out fine, which is embarrassing and informative.

Sample sizes

You will not have enough hires for statistical confidence, and that is fine.

Twenty data points will not prove anything and will show you whether your top scorers are obviously failing, which is the thing worth knowing.

Treat it as a sanity check rather than as research.

The practical summary

Reliability: can two people apply it the same way?

Validity: is it measuring what the job requires?

Both are answerable with an afternoon of work, and answering them tells you more about your process than any vendor's brochure.

What to check

Have two assessors ever scored the same candidate independently?

Do you record assessment scores anywhere you could look at later?

Could you say what each stage of your process is measuring?

And has anybody ever compared scores against how hires actually performed?

The point

Reliability asks whether two people apply it the same way; validity asks whether it measures what the job requires.

Both are answerable in an afternoon.

Underlying all of this

Almost everything in this collection reduces to one discipline: write down what the job requires, assess that thing directly, record the evidence, and look at your own outcomes afterwards. None of it requires buying anything, and organisations that do those four things consistently outperform ones running longer processes built from instruments chosen before the requirements were known.

Independent guidance on skills assessment, selection design and fair hiring practice. External tools are included for practical comparison; evidence from the job remains the basis for decisions.