AI Agents Add 30% Code, Zero Extra Issues: Harvard Study
By @sharedot · · 7 pages
- AI Frontier
- AI Coding Agents
- Engineering Productivity
- Code Review
A Harvard study of 718 firms finds AI coding agents boost code output 30% yet resolve no additional Jira issues, as review becomes the bottleneck.
The 30% code surge that changed nothing
A new Harvard working paper, *Artificial Intelligence in the Firm: Bottlenecks in Software Production*, delivers the first large-scale empirical evidence that AI coding agents have not converted coding speed into shipped software. Researchers Fiona Chen and James Stratton analyzed data from the Jellyfish engineering analytics platform covering 718 firms and more than 725,000 workers between 2021 and 2026. The headline finding is stark: agents like Claude Code, Cursor, and Devin increased code production by 30% — nearly 1,500 extra lines per worker-month — but produced no statistically significant increase in resolved Jira issues or completed epics.
Activity up, outcomes flat
The disconnect shows up in every activity metric: workers are pushing 20% more commits and 23% more pull requests per month, yet the business value measured by resolved issues remains flat. 'People could make a huge number of commits quickly, but code review was still the bottleneck,' one junior engineer told the researchers. Forkast News, which reviewed the paper, frames it as a measurement trap — the industry tracks code volume because it is easy, while ignoring how software actually moves from idea to finished product.
Review queues buckle under the load
The mechanism is the code review process. When developers generate code faster with AI, they push more work into the review queue: review times jumped 49%, adding an average of 3.45 days, the share of pull requests requiring revision nearly doubled, and comments per request rose 35%. Fourteen percent more workers are now pulled into review, burdening senior engineers. 'There is only one bottleneck at a time. If you speed up coding, the bottleneck shifts somewhere else,' one senior engineer said. Forkast News reports that despite 80% of firms using AI-assisted review tools, only about 23% of review comments are actually AI-generated — quality control still falls on humans.
Teams reorganize around AI colleagues
Practitioners are responding by restructuring workflows around the new bottleneck. TechGig reports that teams are treating AI agents as hired colleagues rather than tools — giving each agent an isolated virtual desktop with its own file system, browser, and editor, and running tasks through Kanban boards that prioritize spec review over code review. The logic: mistakes are cheaper to fix in a specification than in thousands of generated lines. GitHub even added a setting in February 2026 to disable or restrict pull requests after an influx of 'AI slop at scale.'
What it means for the adoption narrative
Adoption rates climbed from near zero in late 2024 to over 95% by early 2026, largely without adjusted workflows. The paper's authors put it plainly: 'Realizing the benefits of AI adoption requires complementary investments in downstream review capacity, rather than simply increasing code output.' For builders, the lesson echoes a systems principle — speeding up one stage without unblocking the next just moves the pile of work. The likely next moves are automated testing, better review tooling, and rethinking team structure to absorb the AI-generated surge.