Phase one: AI as a reference
How you know you're here: people paste code into ChatGPT instead of searching Stack Overflow.
It's faster than searching, and it's how almost everyone starts. But the AI never sees your codebase, so every answer is a guess about code it can't read. I was essentially a courier between two browser tabs.
What breaks: nothing dramatic. It just doesn't scale, and it teaches people that AI is a clever search engine.
Phase two: AI in the editor
How you know you're here: everyone has an AI assistant in their editor, each configured differently, and output quality depends on who's at the keyboard.
This is a real step up. The AI can see the code. It was also, in my experience, a zig-zag path to a result: the AI wandered off on tangents, added scope, misunderstood the goal or got stuck in loops. I'd estimate I got to the result I wanted about 20 to 30% faster than hand-coding, but it wasn't very satisfying: lots of waiting around, and dragging the solution back to what I wanted.
What breaks: the team. Each developer has a private way of working with their assistant, nothing is shared, and quality varies from one person to the next.
Phase three: a rules file (CLAUDE.md)
How you know you're here: there's a shared instructions file in the repository, a CLAUDE.md, a .cursorrules or similar, with preferences and a handful of rules. In more mature teams it has grown into a manual, covering process, approval steps, examples of good and bad code, and guidance on judgement calls.
This is the first shared artefact, and it helps. In my case the problems got more serious rather than less. The AI ran database migrations without asking, which is the kind of initiative you admire in nobody. It read secrets it had no business reading. It hallucinated a non-existent API, and occasionally introduced major errors with complete confidence.
My response was to write everything down, and the file grew into a handbook. The content was right but the format was wrong. It was painful to maintain, used up a huge amount of the AI's context window, and couldn't be reused across projects. Worse, it didn't get fully read. I had written a staff handbook for a colleague who skim-reads.
Something else was happening through these first three phases that bothered me more than any single incident. Each phase saved some time, but I was becoming less familiar with my own codebase. That made manual fixes harder, which made me lean on the AI more, which made me less familiar still. My trust in the setup was falling as fast as my dependence on it was rising.
What breaks: trust. More words don't mean more compliance.
Why most teams are stuck at phase three
Phase three feels like governance. There's a file, it's in version control, everyone's AI reads it, and the usage dashboards look healthy. Many AI adoption programmes stop here, because they measure adoption: licences issued, prompts sent, lines suggested.
But a rules file only tells the AI how to behave. It doesn't say what each piece of work should achieve, it can't stop the AI doing something it shouldn't, and it doesn't check what comes back. All of that still lands on the reviewer. So the gains stay personal, and they're partly cancelled out by the extra review effort, the rework and the "where did this come from?" conversations.
The jump out of phase three is organisational, not technical. It needs three shifts:
- From personal assistant to team system. Shared context, roles and standards, instead of everyone doing it their own way.
- From writing code to running the lifecycle. The AI connected to the strategy, priorities and specs, so product and engineering pull in the same direction.
- From reactive to proactive. Agents that don't only wait to be asked, starting with automated review before anything ships.
Phase four: spec-driven development
How you know you're here: no AI work starts without a written spec, with acceptance criteria, that a person has approved.
This was the first big leap. A good spec does two jobs at once. Before the work, it tells the AI exactly what to build and what not to build, so I no longer have to sit and babysit while it writes code. After the work, it gives me a yardstick: I'm not asking "does this look right?", I'm checking it against acceptance criteria we agreed before any code existed.
It also moves the human review to where it's cheapest. Changing a spec costs minutes. Changing code that's already been built on costs days. And because the plan is written down, several agents can work on it in parallel without drifting apart.
What breaks: scale. Specs make each piece of work better, but the AI still needs the right knowledge for each task, and someone still has to check the results.
Phase five: knowledge and roles
How you know you're here: the knowledge is split into focused documents and skills that load on demand, and the AI works in defined roles.
I refactored my giant handbook into a set of focused skills. The AI reads what's relevant to the task in front of it, and nothing else. Context stopped being the bottleneck, and the material became reusable across projects.
That made roles possible: product manager, architect, engineer, QA and release manager, each with its own playbook. An AI reviewing as QA catches things the AI that wrote the code won't. And once there are roles, the work can be run like a team. An orchestrator plans the work and delegates to subagents, each working to a spec in its own role, then checks what comes back before anything moves on.
What breaks: not much, except that everything so far is still advice. A skill can be ignored, and a role can be played badly. Which leads to the second big leap.
Phase six: guardrails and backpressure
How you know you're here: the important rules are enforced by the tooling, and AI output is checked in layers before it reaches human review.
The idea underneath this phase is the difference between probabilistic and deterministic controls. Everything in phases three to five is probabilistic: an instruction, a skill or a role makes good behaviour likely, but never certain, because the AI decides whether to follow it. Phase six adds deterministic controls, which work the same way every time, whatever the AI decides.
This phase has two halves.
Guardrails make the worst mistakes impossible. Instructions are policy, and policy gets ignored under pressure. I learned that the hard way, when an agent ran a delete command with an empty variable and took a couple of dozen applications off my machine. My carefully written rules didn't stop it, because rules are only words. So the hard limits moved into tooling:
- commands run in a sandbox, so they can't reach anything outside the project
- the AI commits through its own identity, with no permission to merge or approve its own work
- secrets and production are blocked outright
- hooks enforce the process: small scripts that run automatically at set points, such as before a command or a commit, and can block it
None of this is novel. It's the principle of least privilege, applied to a new kind of colleague. (Make bad things impossible goes into it properly.)
Backpressure catches the rest. AI produces work faster than people can carefully review it, and review fatigue is how quality dies. So the checking happens in layers, each a filter:
- Automated tests: unit, integration and end-to-end tests, run by the agent while it works and again on every pull request.
- Code quality checks: linting, type checks and shared contracts between the front end and back end, so a change that breaks the other side fails before anyone reviews it.
- Subagents check their own work against the spec's acceptance criteria, and the orchestrator checks it again before accepting it.
- A separate, deliberately adversarial review agent looks at every pull request. It didn't write the code, and it reviews from a different altitude: not "does this work?" but "what's wrong with it?"
- Regular audits sweep the whole app, horizontally (one feature across every concern) and vertically (one concern, such as security or privacy, across every feature).
- A person reviews it last, against the spec they approved at the start, and decides whether it ships.
The first two layers are deterministic, so they catch the same problems every time. The next three are probabilistic, which is why there are several of them. Each layer catches what the one before missed, so by the time work reaches the human review, most problems are already gone. That frees the reviewer to focus on judgement: is this the right thing, done well? And because they approved the spec in the first place, the review is neither too late nor too hard.
What changes: you stop depending on the AI always behaving, and you can start trusting the system instead.
Where trust comes back
Somewhere in phases four to six, the trust problem resolved itself in a way I didn't expect. Trust didn't come back from knowing every line of code. It came back from knowing the system that produced it.
Every CTO already knows this move. Past a certain team size you can't read everything your engineers write, so you trust the checks and balances instead: specs, reviews, tests, release gates. AI doesn't change that discipline. It just gets you there with a much smaller team.
Because Termly holds children's data, this mattered more than usual. AI-written code gets more scrutiny there, not less: nothing ships without a person approving it, and how we build changes nothing about who is responsible for what we build.
A model, not a timeline
To be honest about it, my own route wasn't this tidy. Phases one to three came one after another, but four, five and much of six arrived at roughly the same time, each pushing the others along. Specs made roles useful, roles made layered review possible, and the delete-command incident made guardrails urgent.
So treat the six phases as levels of maturity rather than steps you must take in order. The order still matters as a guide: specs without guardrails leave you exposed, and guardrails without specs leave you checking work against nothing.
Where is your team? A quick self-check
Answer these honestly:
- Is there a shared, version-controlled set of instructions that every developer's AI uses? (No: phase one or two.)
- Does every piece of AI work start from a written spec, with acceptance criteria, that a person approved? (No: phase three.)
- Is your AI's knowledge split by task and role, with work delegated and checked rather than done in one long conversation? (No: phase four.)
- Could the AI merge its own work, read a production secret, or delete something outside the project if it got it wrong? (Yes: you're not at phase six.)
- Is AI output checked in layers, by other agents and regular audits, before a person reviews it? (No: you're not at phase six.)
If you answered "no" to the second question, you're in good company. Most teams are.
Honest limits
This model comes from one team's journey, mine, and from the teams I've compared notes with. Phases blur at the edges, and a team can be at phase five for one kind of work and phase three for another. It's also written from the perspective of a small team. Scaling the same discipline across many developers is a related but harder problem, mainly because it's about habits, not files.

Find out more
If you'd like an outside view of where your team sits, and the right-sized next step, start with the AI delivery assessment. It takes three to five days, and it uses these six phases as the yardstick.
Frequently asked questions
Why isn't AI making our development team faster?
Usually because AI is being used as a personal assistant with a rules file, rather than as a team system. Individuals save time, but without specs, roles and layered review, the savings are lost to rework, inconsistent quality and longer reviews.
What is an AI maturity model for software teams?
A way to describe how a team uses AI, from ad-hoc copy and paste to a governed delivery system. The six phases: AI as a reference, AI in the editor, a rules file, spec-driven development, knowledge and roles, and guardrails and backpressure.
Is a CLAUDE.md file enough to govern AI coding?
No. A rules file tells the AI how to behave, but it doesn't define what each piece of work should achieve, can't enforce its own rules, and doesn't check the output. You need specs, enforced guardrails and layered review as well.
What is spec-driven development with AI?
Writing a short specification with acceptance criteria before any AI work starts, and having a person approve it. It tells the AI what to build, lets several agents work in parallel, and gives you a clear yardstick for judging the output.
What's the difference between probabilistic and deterministic controls for AI agents?
A probabilistic control, such as an instruction in a rules file, makes good behaviour likely but not certain, because the AI decides whether to follow it. A deterministic control, such as a permission, a sandbox, a hook or a test, behaves the same way every time, whatever the AI decides. Use instructions for guidance, and deterministic controls for anything that must never go wrong.



