#contributionLarge real-world assessment dataset
The study analyses 239,521 ratings by 12,649 judges across 193,128 papers and 5,038 journals, giving the variance-partitioning analysis substantial empirical scale. This enables comparisons between judge, paper, journal, and judge-slope components that smaller inter-rater studies cannot make.
↳ Abstract; Results, opening paragraph
#methodological rigourAppropriate variance-partitioning design
The modelling strategy directly matches the research question by estimating paper and judge random intercepts, adding journal effects, and then adding judge-specific random slopes for latent paper dimensions. This lets the paper separate level noise from pattern noise rather than treating all disagreement as residual variation.
↳ Methods, Multilevel models; Results, Unmasked pattern noise
#methodological rigourRobustness checks address key assumptions
The paper tests the main pattern using no-singletons restrictions, richer-classification restrictions, permutation testing, multiple optimisers, and an ordinal Bayesian model. The permutation result, where judge-slope variance falls from 49% to 0.6%, is especially informative for ruling out model-flexibility artefacts.
↳ Results, Findings were robust across model specifications; Figure 1
#impact potentialClear practical audit pathway
The Discussion translates the empirical finding into a concrete proposal: routine noise audits that partition variation between judges and research outputs. This makes the policy relevance legible for funders, journals, and research assessment exercises.
↳ Discussion, implications for research assessment