How Nabu Evaluates Research

Cite as: Nabu Science (2026). Nabu Evaluation Framework v5.9. nabu.science/methodology

Have a methodologist’s eye? We’d like to hear from you. - submit feedback →

Validation Highlights

Validation corpus. Numbers update as the expanded corpus completes.

1. Abstract

Nabu assesses papers on intrinsic merit — Study Quality, Reliability, and Readiness — benchmarked against the standard for field and methodology type, blinded to journal, citations, or author.

Every review leads with Key Claims: 2–5 central assertions in the paper’s own words, each rated on how well the paper’s own evidence supports it. Study Quality comes from four dimensions (Contribution, Methodological Rigour, Reporting, Positioning), scored by multiple blinded AI reviewers and documented evidence. A separate Reliability flag (Red / Amber / Green) surfaces concerns that could change how findings should be read. Readiness places the work on a five-stage ladder from initial observation to real-world use. These three signals never blend — a flag doesn’t lower a score, and a score doesn’t set a claim’s strength.

The rubric is explicit and calibrated, which keeps AI judgments evidence-bound and forces “insufficient information” calls instead of guesses.

Validated four ways on an OECD-wide corpus: Nabu reviews score slightly above the best human reviewers on average (6.1 vs. 5.0; none scored below 4; n=50), hit 0.81 inter-rater reliability (n=400+), and blind-flagged 85%+ of retracted papers via low Study Quality or a Reliability flag (n=50; see §4 for the detection rule). A Nabu review is an independent structured review — not journal peer review, and not a replacement for it.

2. Methodology

Every paper is evaluated on three independent signals - Study Quality, Reliability, and Readiness - and every review leads with the paper’s own Key Claims and how well its evidence supports each of them. All four come out of the same evaluation, with criteria calibrated to the paper’s methodology type and field of science. They are scored, displayed, and reasoned about separately: no flag lowers a score, no score sets a claim’s strength, and nothing is blended into a single number.

2.1 What we score

The sections that follow are in the order a review presents them, so this page can be read alongside a paper page from the top down.

2.2 Key Claims

Every review opens with the paper’s claims, in the paper’s own words, each carrying a rating of how well the paper’s evidence supports it.

Claims are adjudicated from the same reviewer output that produces the dimension scores. No claim rating moves a Study Quality score, and no Study Quality score sets a claim rating.

What counts as a claim

A key claim is an assertion about what this study found, or what the authors conclude follows from it - an empirical result, effect, comparison, or interpretive conclusion drawn from the study’s own evidence.

Wording is the paper’s own. A claim is quoted verbatim or near-verbatim and trimmed to fit, never reworded, and every scope qualifier in the original - dataset, benchmark, cohort, subgroup, condition - is preserved, because dropping one states a broader claim than the paper made.

Primary and secondary

A review carries two to five claims, exactly one of them primary, and the primary is listed first.

The four strength levels

Each claim carries one of four levels, drawn against the evidence the paper itself reports:

LevelWhat it means
Well-supportedthe design directly identifies the claim and the evidence is consistent and adequately powered for it
Supportedthe evidence substantially backs the claim with noted limitations that do not change its direction
Tentativethe claim outruns the design (for example, causal language on observational identification) or rests on evidence with material weaknesses
Unsupportedthe reported evidence does not back the claim as stated

The lineage is GRADE’s certainty-of-evidence rating, which grades how much confidence a body of evidence warrants in an estimate rather than grading the study that produced it. Nabu departs from GRADE in scope: GRADE rates a body of evidence assembled across studies for one outcome, and this rates a single claim against the single paper that makes it. Nothing outside the paper - no replication, no later literature, no meta-analytic pooling - is in the frame.

2.3 Study Quality

The Study Quality score is built from four weighted dimensions. Each dimension is scored holistically against a set of guiding criteria rather than many individually scored sub-components.

  • Contribution · 25%

    Does this paper move the field forward?

  • Methodological Rigour · 45%

    Is the methodology sound for the question asked?

  • Reporting · 10%

    Is there enough detail, clearly presented, to evaluate and reproduce the work?

  • Positioning · 20%

    Does the paper engage fairly with relevant and contradictory prior work, and represent its own novelty and limitations honestly?

Study Quality score calibration

Score rangeWhat the paper looks like
4.4 - 5.0ExemplaryHolds up on every dimension, with no material gaps.
3.8 - 4.3StrongClearly strong overall, with only minor weaknesses on any single dimension.
3.0 - 3.7SoundReal strengths alongside meaningful gaps across one or more dimensions.
2.0 - 2.9LimitedNotable weaknesses across several dimensions limit what the work establishes.
0.0 - 1.9Very LimitedSerious shortfalls on the dimensions that matter most.

Scores are calibrated against published work at the time of publication, not against current standards.

Insufficient Information

A reviewer who cannot extract enough from the paper to judge a component marks it Insufficient Information rather than guessing. II is not a low score. The dimension is dropped from the composite and its weight is redistributed proportionally across the dimensions that could be scored, so a paper is neither rewarded nor penalised for what it did not report - it is simply assessed on less. The review says which dimensions were affected, and a paper carrying several II dimensions is a low-confidence composite that should be read through the dimension breakdown rather than the single number.

Strengths, Limitations, and “In brief”

Strengths and limitations are published with each review. The selection rule is significance: they are the most consequential the reviewers found. Where a Red reliability flag is raised, at least one limitation has to address the concern behind it, so the flag is never left unexplained in prose.

Each is tagged with the dimension it belongs to and carries a locator - the section, table, or figure it came from.

“In brief” is the one-sentence reader summary at the top of the review. It is written by the adjudicator from the completed evaluation - what the paper accomplishes, and the constraint that most limits it - and it is a summary of the assessment, not of the paper’s abstract. It is derived from the same adjudicated output as everything else on the page and adds no judgment of its own.

2.4 Reliability

The Reliability flag is a pre- and post-scoring gate that flags methodological, statistical, or evidentiary concerns.

A Red flag does not mean "bad research." It means "proceed with caution - specific concerns identified." Rigorous work can carry an Amber or Red flag if reliability concerns are present, and weak work can carry a Green flag if no specific concerns are identified.

Reliability does not modify Study Quality. There is no arithmetic anywhere in which the flag lowers a dimension score or the composite - "never blended" is literal, not a figure of speech. A paper reads as a score and a flag, side by side, and the reader does the combining.

  • Internal Coherence

    Methods–results alignment. Statistical results computable. Claims supported by evidence presented in the paper.

  • Research Conduct

    Ethics approval declared. Conflicts of interest disclosed. Preregistration referenced with identifier. Data and code availability as reported.

  • Reference Integrity

    Every cited reference resolves to a real publication whose metadata matches the citation. Topaz et al. 2026, The Lancet

  • Post-Publication Record

    Retraction status. Post-publication concerns raised by the scientific community.

Each component returns one of three statuses: Clean (no concerns identified), Noted (concerns present but do not change interpretation of core findings), or Concern (concerns that could change interpretation if confirmed). The overall Reliability flag is derived from the worst component status across the four.

  • Green — No concerns: No concerns identified. The Study Quality score and the Readiness position reflect merit on their own.

  • Amber — Caveats: Caveats noted. The concerns do not change the interpretation of core findings, but readers should consult the rationale.

  • Red — Concerns: Concerns identified that could change the interpretation of findings. The paper carries an explicit reliability flag in downstream displays; the flag does not lower any Study Quality score.

Confidence

Every review carries a confidence level - high, medium, or low - and individual dimensions can carry their own. Confidence is not a judgment about the paper. It is a statement about how firm the assessment of it is.

What to do with a low one. Read the dimension breakdown rather than the composite, and read the evidence locators rather than the summaries. Low confidence marks the reviews where the underlying reading was genuinely contested, which makes them the ones most worth a human going back to the paper.

2.5 Readiness

Readiness answers one question: where does this paper’s evidence sit on the path from an initial observation to something anyone can use? It is reported as a position on a five-stop ladder.

The ladder

Five stops, weakest to strongest, so a reader who lands on a rung can see how far up it sits and what stands on either side of it.

  1. Foundational

    The work establishes something about how the world is, without yet showing anyone what to do with it. Use is a later question, and often somebody else’s.

    Establishes basic understanding. Application is downstream.

  2. Mechanistic

    The effect has been pinned down under conditions the researchers controlled. It behaves as described in the setting it was studied in; nothing yet says it survives leaving that setting.

    Mechanism characterised under controlled conditions.

  3. Demonstrated

    Someone has built the thing and shown it working once, in a lab or a pilot. It is a demonstration that the idea can be made real, not evidence that it holds up at scale.

    Proof of concept in a lab or pilot setting.

  4. Validated

    It has been tested in conditions close to the ones it would actually meet — real populations, real settings, real messiness — but not yet in day-to-day operation.

    Tested in conditions approximating real-world use.

  5. Deployable

    It has been run in an operational setting and the paper says enough about how to run it that someone else could. This is the top of the ladder, and it is rare.

    Operationally tested, with implementation guidance.

Readiness and Study Quality are separate

A perfectly executed study of a narrow question can sit high on Study Quality and on the bottom rung of Readiness. An ambitious paper aimed squarely at deployment can sit high on Readiness with a low Study Quality score. Neither combination is a contradiction: how well the work is built and how far along the path to use it sits are different questions, asked separately, and the framework never resolves them into a single verdict.

2.6 Role of AI in the Pipeline

Nabu’s reviewers are AI reviewers operating inside the structured rubric, with a human escalation layer above them. Adjudication is also AI-executed within the rubric, with an escalation and quality-assurance path to human reviewers. The system is engineered so that the rubric does the work, not the model.

“The rubric is the evaluator. The AI is the instrument.”

Four operational guardrails constrain the AI to evidence-based judgments:

  • Component-level scoring, not holistic judgment.

    Each reviewer scores each rubric component independently, not the paper as a whole.

  • Evidence-bound rationale per component.

    Every component score must be accompanied by rationale that cites specific text, methods, or results from the paper.

  • The Insufficient Information (II) flag.

    When a reviewer cannot extract enough evidence from the paper to score a component, the reviewer marks the component II rather than guess.

  • Field-specific calibration anchors.

    Each rubric component is calibrated against literature and standards for the paper's Field of Science. This prevents the AI from defaulting to a generic prior of good research.

Two further controls operate inside this pipeline:

  • Bias control

    Reviewers are blinded to author, journal, and institutional information; calibration anchors are field-specific to prevent default-to-prestigious patterns.

  • Anti-metric gaming

    Rubric components reward methodological substance, not surface signals; scores cannot be improved by edits that do not reflect underlying quality.

2.7 Replication Outlook

In developmentA forward-looking estimate of whether a paper’s primary claim would survive independent testing. It is being scored on evaluations now and may appear beside the primary claim on recently evaluated papers. It is a different question from claim strength - which is support inside this paper, not a forecast about the next one - and this section will be published once that distinction, the method, and its calibration can all be documented together.

3. The review as a living document

A review is issued, not closed. Authors can put a note on the record, readers can register where they think the assessment landed wrong, and the framework itself moves. What follows is how each of those works and what each can and cannot change.

3.1 Author Note

An author of the paper can attach one note to its review. The note is published alongside the assessment, in the author’s own words, under the heading “From the Authors”.

Three kinds of note

  • Factual correction

    The evaluation states something the paper does not. Point to the section, table, or figure.

  • Additional context

    Context that shapes how the work should be read but isn't in the paper. Design constraints, ethics limits, preregistration, data availability.

  • Update since publication

    What's changed. Errata, corrected version, replication, follow-up work.

The rules of the record

Every note is moderated before it appears. At most one note is live on a paper at a time, and each author gets one submission per paper, ever - approved, rejected, or withdrawn, that is the one shot. There is no threading: a note is a statement on the record, not the opening of an argument, and the review does not reply to it.

A withdrawn note stays in the permanent record. What was said and when it was said do not disappear because the author later preferred that they had not been said.

Whether a note can change the score

3.2 Field Feedback

Field Feedback is where a reader who knows the area registers their own calibration against Nabu’s. It runs on the paper page beside the assessment, on the three signals that carry a band: Study Quality, Reliability, and Readiness. A respondent picks the band they would have given on each - not a score, the same band vocabulary the review uses - so agreement and disagreement are directly comparable.

Disagreement has to be argued

A respondent who shifts the Reliability band away from Nabu’s must say why, in substantive prose, before the submission is accepted. That is the one dimension where a shift is a claim about the paper’s integrity, and a bare vote is not one. Commentary on Study Quality and Readiness is invited but not required.

Any submission carrying commentary goes to moderation before it is published; a bands-only submission counts toward the distribution immediately. Approved commentary appears publicly beneath the summary, shown against the band the respondent gave and the band Nabu gave, so a reader can see what was being disagreed with.

What the field can set in motion

When three or more respondents shift the same dimension on the same paper, that paper and dimension are raised into an editorial review queue. The threshold surfaces the signal; it does not fire anything. Nothing is re-evaluated automatically on a vote count - a person reads the commentary behind the shift and decides whether the paper goes back through evaluation, and records that decision. Sustained disagreement from people who know the field is a reason to look again, not a mechanism for editing a score by majority.

3.3 What triggers a re-evaluation

A paper returns for a fresh review when something changes that a review of the text should have accounted for: an approved factual correction from an author, an editorial decision taken on sustained field disagreement, a corrected or updated version of the paper itself, or a retrieval failure that meant the original review ran on less of the paper than it should have.

Separately from re-evaluation, the reliability layer keeps moving on its own. Retraction status and post-publication commentary are refreshed against the record after the review is issued, so a paper that is later retracted shows that, whatever it scored when it was read. Those signals reach the reliability flag and the record; by construction they never reach a claim rating.

3.4 Framework versioning

The framework changes. Every evaluation therefore carries the framework version it ran under, printed in its own header, and is citable as issued under that version. This page documents the current version, v5.9.

Two consequences, stated plainly because they are the ones that matter to anyone citing a review. Prior reviews are not silently restated. A published review is not rewritten when the framework moves under it; it keeps the version stamp it was issued with, and a reader can see that the rules it was scored against are not the ones described here. And a re-evaluation produces a new review rather than editing an old one, so the earlier assessment and the later one are both on the record, each under the version that produced it.

4. Validation

Initial validation results from the Nabu evaluation corpus. Numbers will be updated as the expanded validation corpus completes.

Result 1: Critique quality vs. expert human reviews

Nabu’s review critiques were scored on H-Max - a metric that calibrates critique quality against the full set of human expert reviews, where the best human review anchors at 5.0. The approach follows ScholarPeer, a peer-review framework published by Google. Nabu scored 6.1, above the best-human-review anchor (n=50). Notably, none of Nabu’s reviews scored below 4 - none fell short of the human benchmark.

6.1
Nabu critiques
mean H-Max
5.0
Best human review
benchmark anchor
01234567Best human review(anchor)5.0Nabu critiques(mean H-Max)6.1Best-human-review anchor

H-Max calibrates critique quality against the full set of human expert reviews. Approach based on ScholarPeer, a peer-review framework from Google (arXiv 2601.22638).

Result 2: Reviewer agreement

Nabu’s primary reviewers achieve an ICC2 of 0.81 (absolute agreement) across all scoring dimensions (n=400+, sampled across OECD Fields of Science) - more than double the published meta-analytic benchmark for human peer review (0.34, across 48 studies and 19,443 manuscripts).

Bornmann, Mutz & Daniel (2010). A reliability-generalization study of journal peer reviews: a multilevel meta-analysis. PLOS ONE, 5(12), e14331. doi:10.1371/journal.pone.0014331

0.000.250.500.751.00Human peer review(meta-analysis mean)0.34Nabu reviews0.81Good agreement threshold

Result 3: Retraction detection

A curated corpus of confirmed-retracted papers (n=50), evaluated blind against a control set of non-retracted papers from the same sources - with no knowledge of retraction status available to the reviewers.

The detection rule. A paper counts as flagged when it receives a low Study Quality score or a material reliability flag. It is one rule, stated here and used everywhere Nabu quotes this number - the two signals are separate by design, and a screen that ignored either would understate what the framework surfaces.

  • 85%+ of retracted papers were flagged under that rule - by a low Study Quality score, a material reliability flag, or both.
  • The low scores were driven by specific, documented methodological concerns - not a generic “this seems bad” signal.
  • The remainder were papers retracted for reasons not reliably visible in the rubric-scorable text alone (post-hoc data fabrication, image manipulation, ethical violations).

The rubric identified what post-publication scrutiny later confirmed.

5. Limitations

The framework is calibrated against published research. It is most reliable for paper formats where methodology, claims, and evidence are explicit and reportable. It is less reliable, by design, in the following cases:

Reliability and validation results in the Validation section should be read with these scope conditions in mind.

6. Commitments

  • Open methodology.

    The rubric, weights, scoring criteria, and adjudication principles are published in full. Evaluate the evaluator.

  • Full blinding.

    No author, journal, institution, or citation signals enter the evaluation. Publication year is retained solely for era-appropriate calibration. Citation counts and field-weighted citation metrics are displayed on a paper page as post-hoc context: they are never inputs to any score, and they are never visible to reviewers. They are there so a reader can see how the field received a paper next to how the paper reads — which is precisely the comparison the venue-and-citation default makes impossible.

  • No un-paid reviewers.

    All reviewers engaged by Nabu for calibration, escalation and quality control are dedicated professional reviewers trained and compensated accordingly.

  • No conflicts of interest.

    Nabu has no relationship with publishers, journals, or institutions being evaluated. Reviewers have no career incentive tied to the scores they produce. The review is structurally independent: there is no scenario in which scoring a paper higher or lower benefits Nabu commercially.

  • Venue-independent.

    The same rubric everywhere. A paper is a paper.

  • Living reviews.

    Post-publication signals - replications, retractions, method validity - update the reliability layer over time. Field feedback runs alongside the assessment: readers can register their own calibration on Study Quality, Reliability, and Readiness. Where the field meaningfully disagrees with the Nabu assessment, the work is raised for editorial review and can be sent back through evaluation. Authors can put a note on the record. See section 3.

  • DORA and CoARA alignment.

    Paper-level, methodology-based, venue-independent. The framework operationalises what 3,000+ signatory institutions committed to.

What is DORA? → · What is CoARA? →