Quality & measurement in the AI era

The metric you use to judge people is the reason your data lies

· 4 min read

When a metric is used to evaluate people or teams, they start protecting the number instead of the reality it was meant to describe. In AI-assisted teams this shows up as healthy-looking burndown charts built on commitments sized for pre-AI capacity. The fix is to measure the system rather than the people, and to use metrics to ask questions instead of to hand out grades.

A nervous character paints bars to identical heights on a stepladder while a manager gives an enthusiastic thumbs up.

At some point after AI tools landed, the teams I was watching started hitting their burndown. Sprint after sprint, the lines went cleanly to zero. If you only looked at the charts, it was the best period of delivery anyone had ever seen.

It took me a moment to see what was wrong, and it wasn’t the charts. They were accurate. Nobody was faking anything.

The problem was what the charts were measured against. The teams were still committing to the same amount of work they used to commit to before AI changed how fast that work could be done. Execution had improved so much that a sprint’s worth of work could often be closed in a few days. The scope hadn’t moved. So of course everyone hit their targets. The targets were set for a world that no longer existed.

I noticed because I had a rough sense of how long things took now, and the sprint contents looked exactly like they did two years earlier. Same size, same shape.

It’s tempting to call that gaming. I don’t think it was, and the distinction matters. When a number is used to judge you, the rational thing is to protect the number. You commit to what you know you can hit. For years, “what you know you can hit” was close to what you could actually do, so the protection cost very little. With AI, the gap between what a team can hit safely and what it could actually deliver became enormous, and the metric hid all of it.

This is one of the oldest ideas in measurement. The anthropologist Marilyn Strathern put it in one sentence, building on the economist Charles Goodhart: “When a measure becomes a target, it ceases to be a good measure.” I wrote about the same trap in The Art of Agile Metrics, as one of the most common ways metrics go wrong. The example I used then was counting tasks completed per person per sprint. People commit to more than they can deliver, and the quality debt shows up a few sprints later. The mechanism hasn’t changed since. What changed is the size of the distortion. AI multiplied it.

There is a bigger lesson under this, and for me it was the real turning point. Hours of work have almost stopped being a meaningful unit. Burndowns and story points both quietly assume that effort and output move together. When one unit of effort, one person with good tools, produces far more than it did before, measuring effort tells you almost nothing about value.

So what do you measure?

First, stop using metrics to grade individuals or compare teams. I would go as far as the rule I set in the book: avoid individual metrics wherever you can, and if you do collect them, never publish them. The moment a number becomes a grade, people start managing the number, and you lose the one thing you wanted from it, which was the truth.

Second, measure the system. How long does it take for an idea to reach a customer? How often does what we ship actually achieve what we said it would? Those numbers are harder to protect, because they are not owned by any one person, and they keep telling the truth when the speed of execution changes.

And use metrics as questions. When a chart looks too good, the right response is curiosity, not congratulations. A perfect burndown is not a result. It is a prompt to ask what the commitment was based on.

None of this requires a new tool. It requires deciding what the numbers are for.

One thing to try tomorrow: take the metric your team is judged by most often and ask one question about it. If this team became three times faster tomorrow, would this number show it? If the honest answer is no, the metric is measuring how well the team keeps its promises, not what it is capable of. That is worth knowing before your next planning session.

The Art of Agile Metrics covers the common pitfalls of measurement in depth, along with the metrics I would use instead and how to roll them out without breaking trust. Read The Art of Agile Metrics.

Related: Agile isn’t dead. It’s running on assumptions AI already broke. · AI made your team write code faster. Your delivery didn’t notice. · Quality used to mean “no bugs.” It can’t anymore.

Written by

David Tzemach: Engineering Operations Manager and author of REDEFINE, Agile Quality, The Art of Agile Metrics, and Agile Testing Mastery.