Large contemporary real-world grant corpus
The analysis covers 2,267 confidential applications from four 2023 ESRC and EPSRC calls, substantially expanding the scale and contemporary relevance of prior national-grant studies.
↳ Methods §2.1
Assembling the evidence…
Purpose: Assessing grant applications is time-consuming and difficult, adding to the overall burden of academic peer review. Whilst funders are exploring whether AI can help, there is no published research into the accuracy of Large Language Models (LLMs) for scoring contemporary grants. Design/methodology/approach: This study investigates whether six open-weight LLMs (Gemma 3 1B/4B/12B/27B, DeepSeek R1 32B, Qwen 3 32B) can give useful scores for 2267 recent UK Economic and Social Research Council (ESRC), and Engineering and Physical Sciences Research Council (EPSRC) grant applications, comparing them with scores from the original reviewers and funding panel members. Findings: Although the LLM scores are individually inaccurate, when averaged and converted to ranks they correlate positively with expert average scores. The best performing LLM, Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26). Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24, suggesting that it scores are slightly weaker than individual reviewer scores. Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14), which is substantially lower than the inter-panellist correlation (mean rho=0.38). Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state, such as by helping identify the weakest proposals for fast-track desk rejections, to replace one human reviewer, or for triangulation to check for bias.
Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26).
directly estimated across four calls with bootstrapped uncertainty and consistent reporting
Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24.
direct comparison of LLM-reviewer and reviewer-reviewer correlations across all four calls
Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14).
directly reported panel comparisons across four calls with bootstrapped confidence intervals
Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state.
retrospective rank alignment does not test triage safety, fairness, or decision outcomes
Derived from the full evaluation — not a separate score.
Strengths
The analysis covers 2,267 confidential applications from four 2023 ESRC and EPSRC calls, substantially expanding the scale and contemporary relevance of prior national-grant studies.
↳ Methods §2.1
The study justifies Spearman correlation for ranking, reports bootstrapped 95% confidence intervals, and compares LLM alignment with inter-reviewer and inter-panellist agreement.
↳ Methods §2.3; Results §§3.1–3.2; Figure 1
The discussion addresses incomplete proposal text, uncertain human-score ground truth, narrow coverage, model bias, and systemic effects, then retains human responsibility for consequential decisions.
↳ Discussion §4; Conclusions §5
Limitations
Only Vision and Approach were available, although the prompts retained criteria covering team capability, resources, ethics, and other material assessed by humans. This weakens direct comparisons for reviewer-replacement claims.
↳ Methods §2.2
Fast-track rejection, reviewer substitution, and bias checking are proposed without prospective evidence on false negatives, fairness, workload, or funding outcomes.
↳ Abstract, Practical implications; Conclusions §5
The data come from one country, one year, two councils, two broad disciplinary areas, and responsive-mode government calls, limiting transfer to other funding systems and application types.
↳ Methods §2.1; Discussion §4
The study's main empirical contribution is its scale and use of authentic confidential applications across four calls. Its rank-based analysis and uncertainty reporting support the descriptive alignment claims, while the incomplete proposal inputs limit direct comparison with human review. The discussion handles benchmark uncertainty and governance risks substantively, but proposed triage and reviewer-substitution roles extend beyond the retrospective design. Impact potential is therefore strongest as a basis for controlled funder trials rather than immediate deployment.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Sound3.6
Confidence mediumThe study adds contemporary evidence from 2,267 UK grant applications across four calls and compares LLM alignment with both reviewer and panel scores. It meaningfully extends a sparse evidence base, although the advance remains an incremental cross-context benchmark.
The strongest overall configuration... had a moderate mean correlation with reviewer Scores.
Spearman correlations, bootstrapped confidence intervals, multiple models, prompt configurations, and human agreement benchmarks fit the descriptive research questions. Comparability is limited because the LLMs received only Vision and Approach sections, configuration selection used the reported calls, and no prospective decision outcomes were measured.
only the vision and approach sections had been provided for the applications
The research questions, analysis, numerical findings, and distinction between reviewer and panel alignment are consistently presented. The practical discussion is explicitly cautious but moves from retrospective correlations to triage and reviewer-substitution possibilities that were not tested.
the correlations seem too weak to replace expert review at the final panel stage
The paper directly compares its results with prior Swedish fellowship and physics-facility studies and discusses uncertainty in treating human scores as ground truth. Limitations concerning proposal coverage, geography, fields, model choices, bias, and governance are traced into a cautious human-responsibility position.
humans need to take responsibility for important decisions
Caveats4 of 4 checks
The central correlations and sample counts are consistent across the abstract, results, tables, and discussion. A minor reversal of reviewer and panel terminology in the Table 1 caption does not affect the reported analyses.
Ethics approval and ESRC funding are disclosed, and the secure-data restrictions are explained. The funder relationship is a disclosed potential conflict because the study evaluates applications administered within the same funding system and funder personnel commented on an earlier version.
Flags: 2 declared / 5 total
35 of 36 checkable references verified
38 references in manuscript 2 are books, websites or datasets — counted, but not index-checkable
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
The models were evaluated on authentic confidential applications under secure conditions, but the study remains a retrospective agreement benchmark. It does not measure false rejections, fairness, workload savings, or decision quality under actual deployment.
recognise the weakness of the evidence that they give
AI-generated, human-governed. Something look off? Contact us to request a review.