Reliability & peer review

Can AI evaluate research reliably?

Only when it's guided & constrained. Nabu doesn't ask a model for an opinion, it forces several models through a fixed rubric and measures their assessment on strength of evidence.

A single language model asked "is this paper good?" gives an opinion you can't audit or reproduce. Thats why Nabu uses several independent reviewer models score every paper blind against an explicit, field-calibrated rubric. An adjudicator resolves disagreement on the strength of evidence, and cases that fall outside what the models can reliably assess escalate to professional human reviewers. The rubric is the evaluator, the model is the instrument. Tested this way, Nabu's reviewers reach an inter-rater reliability of ICC 0.81 — more than double the ~0.34 benchmark for human peer review (Bornmann, Mutz & Daniel 2010) — and the system is benchmarked against open peer reviews on sample papers (ScholarPeer / Goyal et al. 2026). Every score stays traceable to the line in the paper that earned it, so a human can always see what the system saw and overrule it.

Sources

Frequently asked

Doesn't AI hallucinate?

Constraining models to a rubric, scoring blind, and requiring text-level traceability is designed to catch and surface that (methodology).

Is this just one model's guess?

No — multiple models score independently and an adjudicator resolves disagreement on strength of evidence.

Related

Last updated