Reliability & peer review

What is inter-rater reliability, and why does it matter?

It measures how much independent evaluators agree when scoring the same thing. And when it's low, a couple of reviews are partly luck.

Inter-rater reliability, usually reported as an intraclass correlation (ICC), captures whether two evaluators looking at the same paper reach the same judgement. It matters because evaluation systems compare items scored by different reviewers, and that comparison only holds if reviewers are consistent. When they aren't, rankings carry a large random component, so which papers "win" depends partly on who happened to review them (Pier et al. 2018). Traditional journal peer review scores poorly here: a meta-analysis across dozens of studies put its ICC at about 0.34 — closer to chance than to a reliable instrument. Nabu treats reliability as something to measure, not assume: its independent reviewers reach an ICC of 0.81 across all scoring dimensions (n=400+, sampled across OECD Fields of Science) (methodology). The point of a high number isn't that the system is infallible, it's that the result doesn't hinge on which reviewer(s) you drew.

Sources

Frequently asked

What counts as a good ICC?

Higher is better; ~0.34 is weak (human peer review), while Nabu's reviewers reach 0.81.

Does high agreement mean the scores are correct?

No — it means they're reproducible. Validity is tested separately, against retractions and expert review.

Related

Last updated