#contributionCost-controlled baselines expose hidden tradeoffs
Figure 1 compares complex HumanEval agents against retry, warming, and escalation baselines using both accuracy and dollar cost. Five-run results and uncertainty intervals show that similarly accurate systems can differ sharply in cost.
↳ §§2.2–2.3; Figure 1; Table A1; Figure A1
#methodological rigourDetailed and reproducible computational reporting
The appendices document model versions, prices, prompts, hyperparameters, benchmark variants, seeds, and five-run uncertainty estimates. Appendix I also reports released reproduction code and an interactive cost-adjustment application.
↳ Appendices A.1, B.1, D, E.1, and I
#impact potentialRecommendations target identifiable users
The paper distinguishes the needs of model developers, downstream developers, and benchmark maintainers, then proposes cost reporting, generality-matched holdouts, and standardized evaluation scripts. These recommendations connect empirical findings to concrete evaluation and procurement decisions.
↳ §§4–6; Table 1