I run a squad of AI agents. Here is what the work actually looks like.
An AI agent is a language model placed in a loop: it gets a goal, tools it is allowed to use, and permission to decide its next step based on what just happened. A squad is several agents, each owning a different kind of work. The clearest example is how this website is built: a writing agent, a critic, a builder and an automated quality gate, with one person approving. In my experience the limits are not intelligence but tokens, how precisely the job is described, and how requirements are defined before the squad ever starts.
Ask a chatbot how to archive last year’s files and it will tell you how. Ask an agent, and the files get archived.
That one difference is most of what you need to know. Everything else is detail.
An AI agent is a language model placed in a loop. You give it a goal instead of a question, and tools it is allowed to use: reading files, running code, calling a system, querying a database. Then it decides its own next step, looks at what happened, and decides again. Say you ask it to archive last year’s files. It lists the folder, picks out the files from last year, moves them, notices that two are locked, and comes back to you with a question. It keeps going until the goal is met, or until it reaches something it was told to stop and ask about.
The model underneath can be the same one you chat with. What makes it an agent is the setup around it: the goal, the tools, the memory of what it has already done in this task, and permission to act without waiting for you after every step.
The mental model I find most useful is a contractor. When you hire someone to redo a bathroom, you don’t tell them how to hold a wrench. You describe the result, give them the keys, tell them which walls they must not touch, and ask them to call you before the tiles go in. A new contractor gets a small job first. One who has done good work for you for two years gets more room. Agents should be managed the same way, and for the same reason. Advice can be wrong without breaking anything. Action can’t.
I understood all of this in theory for a long time. What made it real for me was not a single agent. It was moving to a squad.
A squad is several agents, each owning a different kind of work: analysis, coding, testing, design, and more. Each one has its own instructions and its own limits, and the work passes from one to the next the way it would between people on a team.
The clearest example I can give is this website. A squad builds it, and I decide what goes live.
A writing agent drafts each article from my notes and from the chapter of my book it comes from. A second agent, the critic, checks the draft against a written standard: no invented details, every statistic traced to a primary source, no filler. A builder agent writes the code of the site and opens every change as a pull request. And a quality gate, a set of automated checks, decides whether a change is allowed in: does it build, do the links work, is the page fast and accessible. My part is to read the text and say yes or no, to approve changes that touch design or infrastructure, and to do the few things only a person can do, like logging in to an account.
Three moments from the past week show what that looks like from the inside.
The first was a check that passed and tested nothing. The builder set up the gate, it came back green, and only later did we notice that the setting names were written wrong, so none of the rules were running. Nothing failed because nothing was being checked. So I asked for proof: break something on purpose and show the gate catching it. The builder removed one page description, and the gate failed on exactly that line and nothing else. A green light is worth something only if someone has seen it turn red.
The second was quiet drift. An article I had already approved came back with two small changes I never asked for: a sentence turned into a link, and a closing line missing. Each one looked reasonable. Together they meant the published text was no longer the text I had signed off on. The rule now is that an approved article gets compared byte for byte with its source before it goes live.
The third was the agent stopping. When a deployment failed, the error lived in a dashboard the builder had no access to. It did not guess a fix. It told me what it could not see and asked me to paste the log.
That last one is what a contractor with clear limits does, and it is the part that made me trust the rest.
I tried to find the part of my own work the squad could not do, and I could not find one. Everything I used to do by hand, it does. That does not mean I stopped looking. I still want to see the final result with my own eyes, because I am not 100% sure the first or second round will be good enough. What matters is that every prompt and every process gets better than the last.
The limits turned out to have nothing to do with intelligence. The first is tokens, the unit an agent spends every time it reads or writes anything, so every step has a cost. The second is how well I describe the job. When a result was bad, the cause was almost always a vague goal or a missing boundary on my side. The third surprised me most, because it sits before the squad even starts: how requirements are defined. They are still written the old way, for a developer who will ask follow-up questions in the hallway. A squad builds what the words say, and builds it fast. The first pictures for these articles came out as polished objects with no life in them, because the brief said no faces. That was the squad doing exactly what it was told.
Much of what I wrote in The Art of Agile Metrics had stayed theory. With a squad, within days it existed as dashboards open to the whole organization, and it replaced the monitoring systems we already had. The distance between an idea in a book and something running for everyone shrank to a few days.
I am careful about how I say this, because it is easy to hear it as “agents can do everything, so humans are optional.” That is not what I learned. The work of doing moved to the agents. The work of deciding what should be done, and of owning whether the result is right, moved to me. It got heavier, not lighter.
If you remember one sentence from this, make it this one: an agent is only as safe as its scope, and only as good as the description of the job.
One thing to try tomorrow: take one repetitive task you do every week and write it down the way you would brief a contractor. First the result you want, in one sentence. Then the list of things it must never touch. If you can’t write that last list, you’re not ready to hand the task to an agent yet, and now you know exactly what to figure out first.
REDEFINE covers this in two chapters: how agents work under the hood, and how to build a squad of them. Read REDEFINE.
Related: Quality used to mean “no bugs.” It can’t anymore. · Agile isn’t dead. It’s running on assumptions AI already broke. · AI made your team write code faster. Your delivery didn’t notice.