Locked cohort enables direct comparisons
All eight configurations evaluate the same 114 essays, preventing shifting denominators and supporting paired Friedman, McNemar, and Wilcoxon analyses.
↳ Methodology, Analytic design and Statistical analysis
Loading evaluation data…
Large Language Models (LLMs) are becoming increasingly used to support higher education assessment, yet evidence of their capability on reproducing authentic institutional grade-band labels remains limited. There is also limited understanding of how different model families vary in their grading behaviour, calibration, and tendency to produce systematic bias under consistent assessment conditions. This study examines whether LLMs can accurately grade university level student assignments using eight distinct pre-trained LLMs. The approach was tested using 114 human-marked university essays, and found that performance differed significantly across model configurations with some models aligning more closely with the human-grades while others showed clear patterns of undergrading. In some cases, models predicted Pass or Fail despite the essays being human-labelled as Merit or Distinction, showing the risk of systematic grading bias under consistent assessment conditions. Overall, the study indicates that large language models may be useful as supervised assessment-support tools, but they are not yet ready to be used independently for marking higher education student writing. Any future use in higher-stakes assessment would require further fine-tuning, a more directive assessment-specific knowledge base, and sustained human supervision.
performance differed significantly across model configurations with some models aligning more closely with the human-grades while others showed clear patterns of undergrading
paired tests and Tables 2–4 directly show substantial configuration differences and patterned undergrading
In some cases, models predicted Pass or Fail despite the essays being human-labelled as Merit or Distinction, showing the risk of systematic grading bias under consistent assessment conditions
Table 3 directly documents frequent below-reference-band predictions in several GPT-OSS configurations
large language models may be useful as supervised assessment-support tools, but they are not yet ready to be used independently for marking higher education student writing
best-run performance remained modest, but one corpus and two reference bands limit the deployment conclusion
Derived from the full evaluation — not a separate score.
Strengths
All eight configurations evaluate the same 114 essays, preventing shifting denominators and supporting paired Friedman, McNemar, and Wilcoxon analyses.
↳ Methodology, Analytic design and Statistical analysis
Prediction distributions, grade-distance error, mean bias, and Distinction-specific metrics show that similar aggregate accuracy can conceal systematic undergrading.
↳ Results, Tables 2–4 and Figures 2–3
The paper explicitly reframes the task as Merit–Distinction discrimination and explains how the single corpus, common prompt, incomplete reasoning comparison, and absent moderation limit generalisation.
↳ Limitations and future research
Limitations
Each model-temperature cell is represented by one output per essay, so run-to-run variability is not quantified, including at temperature 0.5.
↳ Methodology, Model configurations and Analytic design; Table 1
The compared families differ simultaneously in parameter scale and reported reasoning settings, while reasoning effort is not analysed as an independent factor.
↳ Methodology, Model configurations; Table 1
The reference cohort contains only Merit and Distinction essays from one filtered corpus under one shared prompt, limiting conclusions about full-spectrum grading and other contexts.
↳ Methodology, Data source and essay selection; Limitations and future research
The locked 114-essay cohort and paired analyses provide a sound basis for comparing the eight configurations, while Tables 2–4 make directional undergrading visible rather than relying on aggregate accuracy. The principal methodological deductions arise from single-run generation and from attributing differences to model family when parameter scale and reasoning settings also vary. The paper's literature engagement and limitation handling are stronger than its transferability, because the evaluation covers one filtered corpus, one prompt, and only Merit and Distinction reference labels. The Abstract's broad grading language therefore warrants tighter alignment with the narrower task established in the Methods and Limitations.
Nabu’s assessment, alongside the field’s view.
Are you an author of this paper?
Sound3.4
Confidence highThe study adds a matched cross-family comparison that reports directional bias and label-specific discrimination rather than aggregate agreement alone. The advance remains incremental because it evaluates 114 essays from one filtered corpus and two reference bands.
The empirical task is therefore best understood as Merit–Distinction discrimination conducted within a four-band response space.
The locked cohort and paired nonparametric analyses fit the repeated-comparison question, and the metric suite exposes error direction and grade-band performance. Confidence is limited by one generated output per configuration and by model-family comparisons that also vary parameter scale and reasoning setting.
Reasoning effort was not analysed as an independent experimental factor.
The paper develops a followable chain from cohort construction through agreement, bias, discrimination, and temperature analyses, with limitations repeatedly stated. The Abstract's broad university-grading framing is less precise than the Merit/Distinction-only task established in the main text.
The findings should therefore be interpreted as Merit–Distinction discrimination within a four-band response space.
The discussion engages both favourable and cautionary findings in the recent LLM-grading literature and positions the results as context dependent. The limitations section connects the restricted labels, corpus, prompt, reasoning settings, and lack of new moderation to the permissible interpretation.
Results may differ for other disciplines, writing genres, marking rubrics, or prompt designs.
Lower confidence on Contribution, Methodological Rigour — domain match limited.
Caveats4 of 4 checks
The reported statistics, tables, and matched-cohort analysis are internally consistent. The Abstract nevertheless creates mild scope overstatement by describing general university-assignment grading before the main text narrows the task to Merit–Distinction discrimination.
The paper declares secondary use of de-identified corpus materials, reports no conflicts, and states that derived outputs and the locked-analysis script accompany the supplementary material. No supplied-text conduct concern was identified.
Flags: 4 declared / 5 total
25 of 26 checkable references verified
26 references in manuscript
No retraction notice found in Retraction Watch.
Sources: Retraction Watch ✓
Where this paper’s evidence sits on the path from initial observation to real-world use.
Authentic human-marked essays make the evidence more practice-proximate than a synthetic benchmark. The evaluation remains retrospective and does not test an operational grading or moderation workflow.
This suggests possible utility for coarse second-reading, triage or calibration support, but not for autonomous grade assignment.
AI-generated, human-governed. Something look off? Contact us to request a review.