The day the rules didn't help
An AI agent was tidying up some files. It built a delete command from a variable, and the variable came back empty. A command that should have removed one folder's contents instead worked on a much wider path, and a couple of dozen applications disappeared from my machine.
Nothing in my rules file stopped it. The rules said, in effect, "be careful with destructive commands". The agent wasn't being careless in any way it could have noticed. It was confidently doing what it thought I'd asked, with a variable that happened to be empty. Rules are only words, and words don't intercept a command before it runs.
Nothing important was lost, and everything was recoverable. But it changed how I think about AI agents. The question stopped being "how do I tell the AI what not to do?" and became "how do I make sure it can't?"
Instructions are policy, and policy gets ignored under pressure
This isn't a new idea. It's how security has worked for years. You don't secure a company by writing "please don't access the payroll database" in the staff handbook. You give people the access their job needs and nothing else, and you put controls on the actions that can't be undone.
AI agents need the same treatment, for three reasons:
- They're confident. An agent doesn't hesitate the way a person might before running something irreversible.
- They're fast. By the time you've noticed a mistake, it may have done it a hundred times.
- They forget. Instructions compete for the agent's attention with everything else in its context. The longer the session, the more likely an instruction is to be quietly dropped.
So the rule I now apply is simple. If a mistake would be expensive, irreversible or a breach of trust, it should be impossible, not just discouraged.
What "impossible" looks like in practice
These are the layers I use at Termly. None of them is exotic. They're standard security practice, applied to a new kind of colleague.
1. The AI has its own identity, with deliberately small permissions. The AI commits and opens pull requests through a dedicated bot account, not mine. That account can propose changes, but it can't merge them, approve them, or change the build pipeline. The repository rules enforce it, so it doesn't depend on the agent's good behaviour. There are no long-lived credentials sitting around either: access tokens are created when needed and expire quickly.
2. Nothing reaches the main branches without a person. Every change goes through a pull request, and the protected branches can't be written to directly. A person approves every merge. If an agent is having a bad day, the worst it can do is open a bad pull request.
3. Secrets and production are out of reach. Configuration files that hold secrets are blocked outright: the agent can't read them, and there's no approval route that unlocks them. Production data is off-limits entirely. The agent works in a development environment, with development data.
4. Commands run in a sandbox. The agent's commands run inside an operating-system sandbox. It can write to the project and a temporary folder, and very little else. Network access goes through a filter. If my delete-command story happened today, the command would fail at the edge of the project instead of taking anything else with it.
5. Hooks enforce the process automatically. Small automated checks run at key moments:
- before a commit, they scan for secrets and refuse anything that looks like one
- they block destructive commands that can't be resolved safely
- they stop a session finishing while its pull request is failing
These aren't reminders. They're gates.
What still needs a person
Making bad things impossible doesn't make the system safe on its own. It removes the worst outcomes, so people can spend their attention where it's actually needed:
- Deciding what to build. No guardrail stops an agent building the wrong thing well.
- Approving the plan. Specs are reviewed before any code exists, while changing course is cheap.
- Judging the output. Automated checks and a separate review agent catch most problems, but a person decides what ships.
That's why I think of guardrails and review as two halves of one system. Guardrails make the catastrophic mistakes impossible. Review in layers catches the ordinary ones. People make the decisions. (Review fatigue is how quality dies covers the review half.)
Why this matters more when you hold children's data
Termly holds information about families and children. Privacy, safeguarding and data protection had to be designed in from the start, and whatever process I built had to be one I could defend to the parents using the product.
For us, AI-written code gets more scrutiny, not less. How we build changes nothing about who's responsible for what we build. Guardrails in tooling are a big part of how I can say that honestly.
Where to start
If your team uses AI coding agents today, start with these five questions:
- Do your agents use their own identity, or a developer's account?
- Could an agent merge its own work?
- Could an agent read a production secret or a configuration file holding one?
- Do agent commands run in a sandbox, or with the developer's full permissions?
- Which of your rules exist only as written instructions?
Each "yes" to the second and third questions, and each rule under the fifth, is a bad thing that's currently possible.
Honest limits
This is how one small team does it, with one person directing the agents. A larger team needs the same principles, plus the harder work of agreeing them across many people and tools. Guardrails also have a cost. Every limit is occasionally in the way of something legitimate, and you need a clear, human route for those cases. And none of this replaces a proper security review or a penetration test for the product itself. It protects how you build, not what you've built.

Find out more
If your product was built quickly with AI tools and you're not sure what an agent, or anyone else, could currently do to it, the technical review checks ownership, security and access, and gives you a plan.
Frequently asked questions
How do you secure AI coding agents?
Apply zero trust and least privilege. Give each agent its own identity with only the permissions it needs, block access to secrets and production, run its commands in a sandbox, and require human approval for every merge. Enforce these in tooling, not instructions.
What permissions should an AI coding agent have?
Enough to propose changes, and no more: create branches, commit and open pull requests. It shouldn't be able to merge or approve its own work, change the build pipeline, read secrets, or reach production data.
Are written rules for AI agents enough?
No. Written rules guide behaviour but don't prevent mistakes, and agents can drop instructions in long sessions. Use rules for guidance, and tooling for anything expensive, irreversible or sensitive.
Can AI agents delete files or damage a system?
Yes, if they have the permissions to. A sandbox limits the agent's commands to the project, so a mistake can't reach the rest of the machine or other systems.





