AI made your team write code faster. Your delivery didn't notice.
AI coding tools have sharply increased how much code teams produce, but not how much they ship. CircleCI's analysis of more than 28 million workflows and an NBER study of AI coding agents both found output rising far faster than releases. The bottleneck moved to review and accountability, and teams need quality bars built for AI-sized output.
Somewhere in the last two years, a pull request stopped being a pull request. It used to be a few dozen lines someone wrote by hand and could explain line by line, function by function. Now it can be a thousand lines an AI generated in an afternoon, and the person who “wrote” it is still catching up to what it actually does when someone asks them to defend it in review.
Research backs this up at scale. CircleCI’s analysis of more than 28 million engineering workflows found teams producing more code than ever, without a matching rise in how often that code actually ships. A separate NBER study of AI coding agents found code volume up roughly 741%, while the number of released versions rose only about 30%. Teams are generating faster. Delivery isn’t following at anywhere near the same rate.
I watched this happen up close. A pull request that used to run a few dozen lines became hundreds, sometimes well over a thousand, in a single afternoon, and the review process built for the smaller version didn’t bend to fit the new one. Nobody had signed up to review a thousand-line AI-generated diff with the same line-by-line care they’d give a hand-written one, and pretending the process still worked the same way didn’t make the diff any smaller or the risk in it any lower.
The harder problem wasn’t the size of the code. It was the ask sitting underneath it: move fast, and take full accountability for a process that’s now mostly not yours. Reviewers were asked to own the judgment on output they hadn’t produced, and in a lot of cases couldn’t fully trace back to a clear line of reasoning. That’s a fundamentally different job than the one most senior engineers signed up for. Treating it as the same job, just with a bigger diff, is exactly where teams get stuck, and where quiet resentment starts building even when nobody says it out loud.
There’s a second layer under that, quieter and harder to put a finger on. A senior engineer’s edge was never really “writes code fast.” It was years of pattern recognition: the instinct for where a change would break something three systems away, the ability to smell a bad abstraction before it shipped. Hand a growing share of that pattern-matching to a tool, and the thing that made someone senior starts to feel optional, even when it isn’t. People don’t resist that shift because they dislike the tool. They resist it because nobody sat down and told them what their judgment is actually for now that typing speed stopped being the bottleneck.
The fix isn’t reviewing AI output like it’s still hand-written code, checking it line by line at the old pace. It also isn’t waving it through because there’s too much of it to read carefully. It’s building a monitoring layer for the job that actually exists now: tracking what agents produce, at what quality, against metrics designed for this kind of output, not the ones built for a world where a “big PR” meant three hundred lines. Review time per line stops meaning anything the moment line count triples overnight. What matters now is whether the output holds up against a quality bar the team actually defined on purpose, and whether someone can say, specifically, why it passed.
One thing to try tomorrow: pull your last ten merged PRs and check one number: median PR size before your team leaned on AI generation, versus now. If it’s up 3x or more and your review process hasn’t changed to match, that’s where your risk is sitting right now, whether anyone’s measured it yet or not.
If this is the moment your team is in, REDEFINE walks through exactly this: what breaks first when AI accelerates code output faster than your process can absorb it, and how to rebuild delivery, review and quality around what’s actually happening instead of what used to happen. Read REDEFINE: davidtzemach.com/books/redefine
Related: Agile isn’t dead. It’s running on assumptions AI already broke. · The metric you use to judge people is the reason your data lies