Software · 10 October 2026 · 4 min read

More Code, Not More Software: Harvard Study Reveals the Human Bottleneck in AI Agents

In brief: Research conducted by Harvard University across more than 700 software firms and 300 million development events reveals that deploying AI coding agents increased total lines of code by 30% and commits by 20%. However, actual software delivery—measured by closed issues and wholesale features—remained statistically unchanged. The time saved during initial code drafting was entirely swallowed by downstream code review constraints, longer pull request lifecycles, and increased revision cycles.

by Team Mocchi's

More Code, Not More Software: Harvard Study Reveals the Human Bottleneck in AI Agents

Over the past two years, the technology sector has enthusiastically framed generative models and autonomous agents as the ultimate productivity multiplier for software engineering. The promise that an intelligent assistant can scaffold entire features from high-level prompts prompted breathless predictions of instant delivery cycles and ultra-lean engineering teams. Yet, repository telemetry on the ground tells a much more sobering story: typing speed has very little to do with customer-facing software value.

Quantifying this mismatch with unprecedented scale is a new academic study, the findings of which were broken down by Ars Technica. The research effectively dismantles the assumption that more generated code equals faster product shipping, showing how efficiency gains at the authoring stage get completely bogged down in downstream engineering workflows.

The Harvard Findings: Code Inflation Without Delivery Gains

The research, authored by Harvard University economists Fiona Chen and James Stratton, analyzes an extensive dataset provided by engineering management platform Jellyfish. The authors tracked over 300 million granular work events—including commits, pull requests, and ticket updates—across more than 700,000 engineers at over 700 software development organizations from 2021 through March 2026.

Using a difference-in-differences regression model, Chen and Stratton evaluated what happened after organizations introduced either autocomplete coding assistants or autonomous coding agents capable of writing and submitting patches from natural-language instructions. In terms of raw volume, the effect of AI agents is dramatic: organizations experienced an average 30 percent increase in lines of code written, a 20 percent surge in total commits, and a 23 percent rise in pull requests.

However, when examining completed business deliverables, the researchers found virtually zero impact. The resolution rates for Jira Issues and Epics—the actual units of production-ready software features—did not change in any statistically significant way. Companies did not ship more features, nor did they trim headcount; they merely managed a substantially higher volume of raw code to achieve identical operational outcomes.

The Code Review Bottleneck

The root cause of this paradox lies in what the authors identify as downstream constraints: the human code review process. While an agent can generate several hundred lines of code in seconds, verifying whether that code introduces subtle edge cases, security regressions, or architectural anti-patterns still requires meticulous human scrutiny.

According to the study, the introduction of AI agents significantly lengthened the lifecycle of pull requests. Reviewers leave more inline comments, pull requests are far more likely to require structural revisions, and the time required to merge code balloons. The minutes saved by an individual engineer using an agent to write code are effectively handed over to senior colleagues, who must spend extra hours untangling and reviewing machine-generated output.

Rather than streamlining the development lifecycle, AI agents have shifted the engineering workload from writing syntax to critical auditing, turning experienced developers into full-time quality assurance filters for probabilistic code generators.

The Hidden Cost of the Productivity Mirage

This dynamic underscores an enduring engineering misconception: confusing raw activity with actual progress. In modern software development, code is not an asset to be stockpiled—it is a maintenance liability. Every additional line represents future technical debt, expanded test suites, and ongoing operational overhead.

When autonomous agents generate verbose or redundant implementations, they increase repository entropy. An uncontrolled influx of automated pull requests overwhelms continuous integration pipelines and distracts senior engineers from architectural governance and strategic system design. The Harvard findings prove that without modernizing quality assurance and review architectures, organizations adopting coding agents risk choking their delivery pipelines rather than accelerating them.

Mocchi's take

For companies looking to integrate AI into their software engineering workflows, this study delivers an essential reality check: technology adoption cannot just be about handing out Copilot or agent licenses to individual developers. At Mocchi's, our daily experience shows that the true value of an engineering team lies in deep business domain understanding and sound architectural design—two areas where generative agents still fall short. Measuring productivity through lines of code or commit frequency is an outdated trap; unlocking real value requires investing in comprehensive automated test suites, strict architectural guardrails, and a review culture that rewards simplicity and maintainability over raw generation speed.

Further reading

All articles on the Mocchi's blog