Quality & measurement in the AI era

Quality used to mean "no bugs." It can't anymore.

· 4 min read

When AI multiplies the amount of code a team produces, the old definition of quality, code that passed its tests and its review, can no longer be verified by people reading and testing it. DORA's 2024 research found that AI adoption, while it helps individual productivity, is associated with lower software delivery stability. Quality now has to be built in at the prompt, at the review of the whole output, and in layers of testing that used to be too expensive to run.

A tester inspects a tiny bug in a detective hat on a conveyor belt while an ignored sticky note waves for attention.

For most of my career, “good quality” had a working definition everyone understood. The code passed its tests and a person reviewed it. No bugs, or close enough.

That definition quietly assumed something: that a human could read what was being shipped. Not every line, but enough of it to have an opinion.

That assumption is gone. The amount of code a team writes today does not allow quality to be managed the way it was before. When a team produces far more code than it did two years ago, no review process built for the old volume can keep up, and pretending otherwise just moves the risk somewhere nobody is looking.

The data points the same way. DORA’s 2024 Accelerate State of DevOps research found that AI adoption improves individual productivity and job satisfaction, and at the same time is associated with worse software delivery stability and throughput. People feel faster. The system gets shakier.

The popular answer is “AI will test it.” I think that is partly right, and only if it is done properly. An AI will test what you ask it to test. If you ask it to confirm the code works, it will usually find a way to confirm it. Tests written by the same process that wrote the code tend to share its blind spots, so a green suite can mean the code and the tests agree with each other, which is not the same as the code being right.

What changes is where quality lives.

It starts at the prompt. If the description of what we want is vague, the AI will build the wrong thing correctly, with clean code and passing tests. So the first quality check happens before any code exists: is the intent precise enough that someone could tell a right result from a wrong one?

Then it moves to the output as a whole. Instead of reviewing a thousand lines one by one, the question becomes whether the entire change does what was intended, and whether anyone can explain why. If nobody can say why it passed, it didn’t really pass.

And then there are testing layers most teams could not afford before. Negative tests, adversarial inputs, contract checks between services, checks that run continuously against production behavior. Writing those used to cost more than the team had. Now they cost a prompt and some review.

I’ll be honest about how this tends to go, because it is not a straight line. I have seen what happens when a lifecycle goes AI-first. The first move is to shift almost all testing into the AI layer, away from dedicated testers. It looks efficient. Then a small group comes back. Not to test features, but to check that the whole package holds up to the standard the team set. That turns out to be a job AI cannot yet do on its own.

I don’t think this is the final shape either. That small group is a stage, not an end state, and it will last until quality is something the system itself can be trusted to hold. We are not there. Anyone who tells you they are is probably measuring coverage and calling it quality.

When I wrote Agile Testing Mastery, the argument was that testing is a continuous practice the whole team owns, not a phase at the end. AI didn’t make that idea outdated. It made it mandatory, because a quality gate at the end of the line cannot keep up with the volume coming down it.

So here is the definition I use now. Quality means the result does what we meant, holds up in conditions we didn’t plan for, and someone can explain why it passed. Bugs still matter, of course. They just stopped being the whole story.

One thing to try tomorrow: before the next feature goes to an AI, write down three ways it could be wrong. Not how it should work, how it could fail. Then ask the AI to generate tests from that list instead of from the code it wrote. Compare what those tests catch to what your usual suite catches.

REDEFINE lays out this new definition in full, as five layers of quality from intent to outcome, and what a Definition of Done looks like when AI writes the code. Read REDEFINE.

Related: AI made your team write code faster. Your delivery didn’t notice. · I run a squad of AI agents. Here is what the work actually looks like. · The metric you use to judge people is the reason your data lies

Written by

David Tzemach: Engineering Operations Manager and author of REDEFINE, Agile Quality, The Art of Agile Metrics, and Agile Testing Mastery.