Reliability & peer review

How reliable is peer review?

Inconsistent - independent reviewers agree only weakly, and for the strongest submissions their rankings are close to random.

Peer review is essential to get suggestions to improve manuscripts, but as a measurement it is noisy. A meta-analysis across dozens of studies put the inter-rater reliability of journal peer review at about 0.34 - far below what you'd accept from a reliable instrument (min. 0.7) (Bornmann, Mutz & Daniel 2010). When researchers replicated the NIH process and had different reviewers score the same grant applications, agreement was low, and the ranking of the best applications carried a large random component (Pier et al. 2018). None of this means individual reviewers are careless, it means a single review is partly luck of the draw, and one pass can't be treated as a definitive quality verdict. That's the case for structured, repeatable evaluation with measured agreement, rather than one opinion behind closed doors.

Sources

Frequently asked

Does adding more reviewers fix it?

It helps somewhat, but panel discussion has been shown not to reliably raise agreement. What increases the agreement rate is a common rubric with shared definitions and calibration. (Pier et al. 2018).

What counts as a good reliability score?

Higher is better; human peer review sits near 0.34, while Nabu's reviewers reach 0.81.

Related

Last updated