#contributionLarge multi-generation tool comparison
The study compares autocomplete, sync agents, and async agents in a single empirical framework using more than 100,000 GitHub developers. That scope lets the article address whether productivity effects grow across tool generations rather than measuring a single tool snapshot.
↳ Abstract; Introduction; §5 Data and Summary Statistics
#methodological rigourDesign targets AI adoption contamination
The matched event-study design uses control outcomes from the same calendar week one year earlier to reduce contamination from unobserved private AI use among contemporaneous non-adopters. The article also reports activity matching, flat pre-trends, placebo tests with non-AI tools, and comparison to prior field-experimental autocomplete estimates.
↳ §6.1 Empirical Strategy; Figure 5; Table 3
#impact potentialConnects task gains to final output
The article does not stop at commits; it traces attenuation through files, pull requests, repositories, releases, and four app marketplaces. This makes the weak-link claim empirically meaningful by showing that large coding-activity gains correspond to smaller shipped-output and usage changes.
↳ Figure 1; Introduction marketplace summary; §4 Model of Software Production