#methodological rigourInference computed over model replicates
The Welch tests and TOST compare AUROCs across five model replicates applied to the same test cohort (df 4.878), so the p-values reflect variation between models rather than participant-level sampling uncertainty. No participant-level bootstrap or DeLong intervals are reported, and the roughly two dozen hourly tests carry no multiplicity correction.
↳ Methods, Model Application and Evaluation; Results, Discrimination Performance
#methodological rigourCausal attribution the design cannot isolate
The Conclusion states that delayed diagnosis recording in UK Biobank contributes to disease-stage misclassification and inflated performance, but the diagnosis lag is never quantified, stratified or manipulated. Cohort, device (AX3 versus AX6), preprocessing and phenotype (iRBD versus early clinical Parkinson's) differences are fully confounded with the proposed mechanism.
↳ Conclusion; Discussion, Diagnosis Timing; Methods, Preprocessing
#positioningStated power limit contradicted by null conclusions
The limitations section states that modest sample sizes limit statistical power, yet the Conclusion asserts the German cohort showed 'no discernible prodromal signal' and the Discussion describes 'no meaningful model performance', against a reported AUROC confidence interval of 0.56-0.62 that excludes chance. A failure to detect a difference under acknowledged low power does not support a null conclusion.
↳ Limitations and Future Work; Discussion, Cultural Factors; Conclusion
#methodological rigourEquivalence margin makes non-equivalence uninformative
The TOST bound of δ = 0.5·SD corresponds to roughly 0.015 AUROC and is not justified. With five model replicates, equivalence is close to unattainable by construction, so the abstract's statement that no cohort pair demonstrated equivalence carries little information.
↳ Methods, Model Application and Evaluation; Results, Discrimination Performance; Abstract