Reliability

Reliability is the consistency of a measurement — whether it produces the same result when nothing about what's being measured has actually changed.

Reliability is the consistency of a measurement — whether it produces the same result when nothing about what's being measured has actually changed.

A bathroom scale is reliable if it gives you the same weight when you step on it twice in a row. It's unreliable if it swings by ten pounds each time with nothing else different. The same idea applies to psychological measurement: if you took a personality test today and again next week, a reliable test should give you very similar results, assuming your personality hasn't meaningfully changed in the meantime.

Reliability is necessary but not sufficient for a good measure. A scale that's stuck five pounds too heavy is perfectly reliable — it gives the same wrong answer every time — but it isn't valid. That's why psychometricians treat reliability as a floor to clear, not a finish line: an unreliable measure can't be valid, but a reliable one still might not be measuring the right thing.

There are several ways to check it: test-retest (same people, two time points), internal consistency (do items meant to measure the same thing agree with each other), and inter-rater (do two observers rating the same behavior agree). Each catches a different kind of inconsistency.

See also

  • Validity Validity is whether a test actually measures the thing it claims to measure, rather than something else entirely.
  • Classical Test Theory Classical test theory is a foundational framework in psychometrics holding that any observed test score is made up of a person's true score plus random measurement error.
  • Item Response Theory Item response theory is a modern approach to psychometrics that models how the probability of a specific answer relates to a person's underlying trait level, item by item.
  • Construct A construct is a theoretical concept, like intelligence or extraversion, that can't be observed directly but is inferred from patterns in observable behavior.