#contributionMultiple variance sources tested systematically
The experiments examine network architecture, reward scale, random seeds, environment selection, and implementation choice rather than attributing reproducibility problems to one source. Figures 2–6 and Tables 1–3 show that each can materially affect reported performance.
↳ Experimental Analysis, Figures 2–6 and Tables 1–3
#methodological rigourExperimental settings are extensively documented
The supplement specifies network structures, optimization settings, batch sizes, seeds, environments, and implementation modifications. Results are generally accompanied by learning curves, standard errors, or bootstrap intervals.
↳ Supplemental Experimental Setup; Tables 4–14
#impact potentialRecommendations connect directly to experiments
The conclusion translates observed variability into concrete recommendations on reporting hyperparameters, implementations, trial counts, random seeds, and uncertainty. It also identifies RL-specific significance testing as an open methodological problem.
↳ Discussion and Conclusion