A developer finishes a change faster with AI. That is a useful gain. The change still needs review, tests, integration and maintenance. If those stages accumulate a queue, faster authoring does not automatically produce faster delivery.
Known: output and review load can rise together
A preprint submitted on July 2, 2026 followed 802 developers and 196,212 pull requests from January 2024 through April 2026 at a mid-sized, AI-forward company. The company sought to double merged pull requests per engineer. By April, that metric reached 2.09 times its pre-mandate baseline.
Per-reviewer load roughly doubled, automated review overtook human review, and merge and revert rates stayed broadly steady. Gains were concentrated in newer code. Adoption was not randomized: this is evidence from a favorable setting, not proof that AI doubles business value everywhere. Read the primary study →
Known: measuring productivity has selection problems
In its February 24, 2026 update, METR explained why its follow-up developer experiment could not reliably estimate current AI gains. Developers who wanted to keep using AI were less willing to participate in work that might prohibit it. Some withheld tasks they expected AI to accelerate. Lower participant pay also affected selection, while concurrent agents made time accounting difficult.
METR considered larger speedups plausible, but described its data as weak evidence for their size. That is a measurement limitation, not evidence that the benefit is zero. Read METR’s update →
Uncertain: more merged code is not a standard measure of value
Pull requests vary in size, difficulty and purpose. Stable revert rates do not establish long-term maintainability, security or defect costs. The company study neither shows that AI inevitably lowers quality nor establishes a general effect on QA staffing. Broader evidence needs to follow changes through production and maintenance.
A rising merge count can coexist with useful progress, extra rework, or both. Teams need release outcomes alongside output counts to tell the difference.
Speculation: verification may become a larger share of work
If authoring gets cheaper faster than verification, teams may spend more effort defining expected behavior, selecting tests and deciding whether a change is safe to ship. This is a conditional scenario, not a forecast of guaranteed job growth. AI review and testing could absorb some workload; measure them by consequential errors caught and missed, as well as noise and human effort.
| Stage | Question to answer |
|---|---|
| Author | Does the change solve the intended problem, and can its owner explain it? |
| Review | How much waiting, human effort and rework does approval require? |
| Test & integrate | Do critical behaviors and dependencies still work? |
| Deploy & maintain | What happens to defects, incidents and later change costs? |
Signals to watch
- Review queue size and time from review-ready to deployment.
- Human review effort and rework per change, accounting for size and risk.
- Production defects and incidents after release.
- Whether gains persist in mature repositories.
- Whether automated review catches consequential errors without excessive noise.
How to prepare
Developers
Keep changes small, understand accepted code, and supply intent, assumptions and evidence. Make the reviewer’s job easier.
QA teams
Prioritize critical behavior and integration risks. Use tests whose expected results come from requirements rather than simply copying the implementation. Track escaped defects and rework.
Leaders
Compare the whole workflow with a baseline: lead time, review effort and release outcomes alongside output. Check whether saved authoring time survives downstream.
Our read
The useful question is how much reliable, valuable software a team can deliver with its available review and testing capacity. The case study reports substantial output gains in favorable conditions and more work moving downstream. Our practical inference: strengthen verification and measure the whole delivery path.
Primary sources
- He et al. — AI Writes Faster Than Humans Can Review (July 2, 2026, v1 preprint)
- METR — We are Changing our Developer Productivity Experiment Design (February 24, 2026)