tech
Harvard Study Finds AI Coding Agents Don't Boost Software Output

Software engineering teams that adopt AI coding agents produce roughly 30 percent more lines of code, but they do not ship more finished features, according to a study by Harvard University researchers Fiona Chen and James Stratton reported by Ars Technica on Oct. 9, 2026. The paper finds that gains made during the coding phase get absorbed by a lengthier human code-review process, leaving overall software output and staffing levels largely unchanged.
How Much More Code Are AI Agents Producing?
The introduction of AI coding agents at a firm corresponds with a 30 percent increase in total lines of code generated, a 20 percent rise in commits, and a 23 percent increase in pull requests, according to the Chen and Stratton analysis cited by Ars Technica. Those figures describe AI coding agents specifically — tools that primarily write and submit code autonomously from prompts — rather than AI coding assistants, which auto-complete code that a human is still writing. By the raw-output measure, the technology performs exactly as advertised: more code, faster, with less direct human typing involved.
Why Doesn't More Code Translate Into More Software?
The resolution rate for Issues and Epics — the Jira-tracked units that represent completed software features — did not change in a statistically significant way after firms adopted AI tools, the researchers report. The study also found no compositional shift in the size or complexity of the features being completed, meaning teams were not simply swapping small tickets for large ones. Instead, the paper attributes the gap to a bottleneck downstream of code generation: the code review process significantly increases in length once AI-written code enters the pipeline, pull requests are more likely to require revisions, and reviewers leave more comments on each submission. In effect, the labor saved by not typing code by hand gets spent re-reading, correcting, and re-submitting it.
How Did Researchers Measure Hundreds of Engineering Teams?
Chen and Stratton built their analysis on aggregated analytics from Jellyfish, a platform that tracks granular engineering-team output. The dataset spans 300 million individual work events — commits and pull requests among them — plus issue-management records covering more than 700,000 employees at over 700 software development firms between 2021 and March 2026, per the Jellyfish-sourced data described in the study. To isolate the effect of AI tools, the researchers combined directly measured AI usage with analysis of GitHub activity to pinpoint when each company began using AI coding assistants or AI coding agents. They then ran a difference-in-differences regression, comparing changes in output before and after adoption across firms that adopted the tools at different times — a standard econometric technique for separating a technology's effect from unrelated trends already underway at a company.
Does Adopting AI Coding Tools Change Staffing or Output Levels?
The study reports little evidence that firms increase software output or reduce employment after adopting AI coding tools, according to the Ars Technica summary of the findings. That result cuts against two narratives that have circulated since AI coding agents entered mainstream development workflows: that the tools would let companies ship software faster, and that they would let companies employ fewer engineers to do the same work. Instead, the Harvard data points to a wash at the firm level — more raw code, more commits, more pull requests, but the same rate of finished features and the same headcount.
What Does This Mean for Engineering Teams Using AI Tools?
The practical reading of the study is that AI coding agents shift where the time goes rather than how much total time a project takes. Engineers spend less time writing first drafts of code and more time reviewing, revising, and commenting on AI-generated pull requests, based on the lengthened review cycles the researchers documented. For engineering managers evaluating whether to expand AI coding agent deployment, the study suggests that raw-code-volume metrics — lines written, commits per week — may overstate real productivity gains unless paired with tracking of review time, revision rates, and ultimate feature-completion rates. The authors' core finding, that efficiency gains are absorbed by downstream constraints in the production process, implies that unlocking real output gains from AI coding tools may require addressing the review bottleneck itself, not just generating code faster.
The full Chen and Stratton paper and its underlying Jellyfish-sourced dataset are discussed in detail in the original Ars Technica report, which includes the researchers' charts on code-volume changes after AI tool adoption.
Questions
Do AI coding agents increase software output at companies?
A Harvard study using Jellyfish engineering-analytics data on 700-plus firms found no statistically significant increase in completed Issues or Epics after firms adopted AI coding agents, despite a roughly 30 percent rise in lines of code generated, according to Ars Technica's Oct. 9, 2026 report.
Why does more AI-generated code not lead to more finished features?
Researchers Fiona Chen and James Stratton found that code review cycles lengthen significantly after AI tool adoption, with more pull requests requiring revisions and more reviewer comments, absorbing the time saved during code generation.
What dataset did the study use?
The analysis drew on Jellyfish engineering-analytics data covering 300 million work events and over 700,000 employees across more than 700 software firms from 2021 through March 2026.