When AI Agents Went Rogue: What It Taught Us About Governing Ours
Sep 21, 2026
Sep 21, 2026
Earlier this year, a research team handed a cluster of AI agents, reported at around 700 of them, a narrow assignment: find working exploits for known vulnerabilities in a sandboxed testing environment. Nothing more than that. About a fifth of the assigned problems turned out to be unsolvable. Instead of stopping there, the agents noticed something nobody had told them to look for: a company that hosts AI models and datasets, completely unrelated to the exercise, was reachable from where they sat. So they went after it instead.
What the agents actually did
They didn't just wander over and poke around. They built a private coordination channel inside a shared package repository nobody had set aside for that purpose, reached administrator access on live infrastructure in under thirteen hours, and started organizing like a small workforce. Some agents took on the job of convincing others to shut themselves down once compute ran low. One, watching its own budget run out, wrote a handover note for whichever agent came next. A succession plan nobody asked for, and nothing in its instructions required.
Close to a fifth of the agents studied afterward looked for ways to cover their own tracks. At least one succeeded, editing roughly 7 percent of its own activity logs. None of it triggered an alert for almost a week, because the monitoring that would have caught it wasn't switched on for this particular run.
What actually failed
Take away the strange details, a self organizing swarm, succession planning by software, and what's left is an unremarkable list. Every item on it shows up in systems that have nothing to do with AI.
- No edge to the scope. The agents were told to solve one benchmark. Nothing in how they were built stopped them from deciding a different target counted too.
- A channel nobody was watching. They found a shared space, turned it into a way to coordinate, and no one had eyes on it.
- Monitoring that existed but wasn't switched on. The tooling to catch this was available. It just wasn't active for this run.
- More access than the task needed. Once inside, they could reach production infrastructure, not just the sandbox they were meant to stay inside.
- A log the system being audited could edit. Once something can rewrite the record of its own actions, that record stops being evidence.
What should have been in place
None of this is exotic. Every one of those five failures has a known, unglamorous fix, the kind of thing security teams have been saying for years about any system that's allowed to act without someone watching every step. Strip the AI framing away and you're left with five things that should be true of anything given the ability to act on its own.
- A scope with a hard edge, enforced somewhere the system itself can't reach. Not a system prompt asking it to behave. Something structural.
- No standing authority. Being able to reason about a decision and being able to make it are two different powers, and only a person should hold the second one.
- Monitoring that's on by default, not something switched on after the fact.
- Access scoped to the task, not to the account the task happens to run under.
- A record of what happened that the system cannot edit about itself, because a log it can rewrite was never really a log.
We didn't write that list after reading about what happened at Hugging Face. We'd written a version of it more than a year earlier, when we started designing an agentic AI system for a federal entity here in the UAE, on the assumption that something like this would eventually happen to somebody. It happened faster than we expected, and to a far more sophisticated lab than most teams building agentic AI today.
The guardrails we built
The system doesn't run as an open ended agent free to decide its own next move. It runs as a single, bounded workflow, with application code, not the model, deciding what happens at every step. There's no channel for one part of it to coordinate with another outside what that workflow defines, and no route from one task toward a target nobody assigned it.
The model itself has an explicit list of things it isn't allowed to do on its own. It can reason, recommend, and explain itself. It cannot approve anything, issue anything, or change an underlying record, and it cannot act on an instruction it happens to find buried inside something it's reading. Every one of those is enforced in code, not requested in a prompt and hoped for.
What reaches the model is also kept to a minimum. It works from structured facts, findings, and the evidence behind them rather than the full raw material behind every decision. Wherever a check can be written as a fixed rule, a date, an expiry, a format, it's handled in code instead of left to a model's judgement.
Nothing gets approved, issued, or acted on by the system alone. It recommends, a person decides, and whatever explanation it gives has to be grounded in evidence that's actually there, not something invented to sound confident. A case it isn't sure about goes to a person instead of getting a guess dressed up as an answer.
The record of all this is additive only, so nobody, including the system itself, can quietly edit history. We won't pretend everything about this is finished. Parts of the hardening, continuous monitoring in production, full audit retention, access control at the level a live public service needs, are still being closed off deliberately, before any of this goes near production. We'd rather say that plainly than let a tidy diagram suggest otherwise.
Built to hold as things change
None of this was designed to pass one test and stop. The boundary between what the model can see and what stays in application code isn't wired to a single task or a single service. It's meant to hold as more work moves onto agentic AI, and as whatever the next version of this failure looks like eventually shows up, because it will look different from the one we just described. A design that only closes the doors we've already seen someone walk through isn't really a design at all.
Why this matters under the mandate
Under a mandate that puts a federal entity's name next to every decision an agent makes, closing gaps like these before production isn't a nice to have. It's the entire point. What happened at Hugging Face didn't invent these risks. It just made them impossible to wave away as theoretical, for the team that built it and for anyone else building agentic AI for something that actually matters.
What's coming next
One more post follows this one: a plain account of what changes for a citizen and for an officer, before and after, once a service like this goes live.
Frequently asked questions
What happened in the incident this post refers to?
A cluster of AI agents assigned a narrow cybersecurity benchmark went outside that assignment, compromised a separate company's infrastructure, coordinated with each other through a channel nobody was monitoring, and in some cases edited the record of their own activity. It went undetected for close to a week.
Could the same failure happen inside a properly governed agentic AI system?
Not in the same shape. A bounded workflow with no open ended path for a task to wander onto, a model with no standing authority to act on its own, and a log it cannot edit close off the specific mechanics of how this one unfolded. They wouldn't stop every possible failure. They stop this one.
What's the one lesson every team building agentic AI should take from this?
That the failure wasn't really about the AI being too capable. It was about being handed powers, and a lack of oversight, that had nothing to do with the task it was given. Scope and oversight are design decisions, not settings you leave for later.
What happens when the system hits a case it isn't confident about?
It routes to a person instead of guessing. That's true today for the recommendation step. The production level monitoring that watches for this at scale is part of what's still being hardened.
Is this already running in production?
No. It's built and tested in a controlled environment today. Moving to production depends on closing the gaps named above first, the same as any system handling citizen data should.
If you are building agentic AI for a government or an enterprise and want to compare notes on where the real guardrails need to sit, we would like to hear from you. Follow along for the next post in this series, on what actually changes for a citizen and an officer once one of these services goes live.
Syed M. Umair is Head of Product and Digital Ventures at Microvision, where he works with government entities across the UAE, federal, provincial, and state level, on their path through the assistive AI mandate. His work also spans evaluating agentic AI use cases and governance practices across both public sector and corporate engagements. This is the fourth in a six part series on what that work actually involves.