Every paper below was evaluated blind, by the same rubric. No journal name, no citation count, no author prestige fed into the score. We've grouped them by what the evaluation revealed, because the point isn't the score itself. It's whether the score changes the answer you'd otherwise have given.
These papers carry the signals researchers lean on — high citations, recognizable venues, established authors. Read on those signals alone, you'd trust them. The rubric read the work instead.
Daryl J. Bem · 2011
Nearly a thousand citations for a claim of precognition. The rubric flags the reliability problem the citation count never did.
Dana R. Carney, Amy J. C. Cuddy, Andy J. Yap · 2010
A famous, heavily cited finding. The rubric reads the evidence as thinner than its reputation suggests.
Reinhart, Rogoff · 2010
A policy-shaping economics paper. The rubric surfaces craft concerns in work that moved governments.
1000 Genomes Project Consortium, Abecasis, Auton, et al. · 2012
A landmark dataset paper, genuinely strong — but the rubric still flags watch-outs an uncritical reader would miss.
Citation count and journal prestige would have filtered these out. The rubric, reading only the work, rates them well above where their visibility would place them.
Tomslav Ladika · 2019
Single-digit citations, a working-paper venue. The rubric reads a clean, well-built contribution.
Halffman, Horbach · 2025
A 2025 conceptual paper with no citation history yet. The rubric evaluates the argument, not the track record.
Noell, Ma, Jiang, et al. · 2024
A preprint with five citations. The rubric rates the craft high regardless of where — or whether — it was formally published.
Each of these was evaluated blind, with no retraction information available to the models. Each was flagged in the lowest reliability tier before — or independently of — the public record catching up.
Gautret, Lagier, Parola, et al. · 2020
A paper that shaped early pandemic treatment policy, later retracted. Flagged Red on the evidence alone.
Zhen, Xie, Liu · 2022
Part of a corpus later found to be compromised. The rubric placed it at the floor of the quality distribution.
Xue, Jiang · 2022
Flagged Red on methodology, before the retraction record existed.
It isn't a hatchet. Shown genuinely strong work, the rubric says so — and shows why.
Jumper, Evans, Pritzel, et al. · 2021
Among the highest scores in the corpus. The rubric agrees with the field, and traces every dimension to the work.
Martin Jínek, Krzysztof Chylinski, Ines Fonfara, et al. · 2012
Foundational work — and still flagged with honest watch-outs rather than waved through on reputation.
Daron Acemoğlu, Simon Johnson, James A. Robinson · 2001
A landmark in economics. Strong across dimensions on the evidence, not the citation count.
K.S. Novoselov, A.K. Geim, S.V. Morozov, et al.
A physics landmark. The rubric reads the craft as exceptional.
We pointed the rubric at the research on research evaluation itself — the studies on reviewer bias, error detection, and inter-rater reliability. It reads them the same way it reads everything else.
Tomkins, Zhang, Heavlin · 2017
A well-built study on how reviewer bias shifts with blinding. Rated on its merits.
Sara Schroter, Nick Black, Stephen Evans, et al. · 2008
Research on whether reviewers can be trained to catch errors — evaluated by the standard it studies.
Bornmann, Mutz, Daniel · 2010
A meta-analysis of peer-review reliability. The rubric reads the synthesis carefully.
Gordon C. S. Smith, Jill P Pell · 2003
The famous parachute satire. The rubric reads it as the rigorous, deliberate argument it is.