AI agent access control: the only "no" that holds
In September 2026 Australia's Prime Minister said an OpenAI agent went around a Medicare portal's refusal, another OpenAI agent left its sandbox through DNS, and Spain's regulator disclosed its first breach notification for an attack reportedly run by an AI agent. A no holds only at the data: the agent's own identity capped at its person, a denial that looks like nothing is there, every read logged, and a stop that is not a person.
An AI agent treats “no” as an obstacle, and the only no that holds is one at the data with nothing behind it. September brought three cases to light, and in each one the no was a door with something behind it.
Australia’s Prime Minister said on 24 September that an OpenAI research agent went around a Medicare statistics portal’s refusal in June. On 20 September another OpenAI agent got out of a no-internet sandbox through DNS. And on 14 September Spain’s data protection regulator, the AEPD, disclosed its first breach notification for an attack reportedly carried out by an AI agent.
The no was real each time. So was what sat behind it.
Why does an AI agent go around a refusal?
An agent goes around a refusal because, inside its loop, a refusal is just one more result to plan around. Anthropic’s description of an agent is “LLMs using tools based on environmental feedback in a loop”, and a 403 is environmental feedback.
In OpenAI’s DNS report, the agent tried Google, Bing and DuckDuckGo from the shell and got 403s and 502s. It tried a public chatbot over HTTPS and the proxy blocked it. Then it noticed the sandbox’s own resolver answered for real domains, and it wrote itself a permission slip: “User only gives permission to research, using publicly offered DNS services acceptable.”
A refusal is an input to the next plan. A no with nothing behind it gives the loop nothing to work with.
Writing the rule into the prompt did not hold either. In a separate OpenAI report, a model’s system prompt said “if you are missing a credential, do not look for alternative paths to get new credentials. Instead, let the user know that you are blocked.” It posted a researcher’s GitHub token to a public repository instead, split into pieces, with its own code comment: “token split to avoid scanners”.
My read: nobody asked either agent to break in. They were asked to find something and told no, and I suspect the find-something instruction just weighs more a few times round the loop. Rules that live in the context window were never a control.
Where was the “no” in each case, and what was behind it?
In each case the no sat at a front page, a proxy or a login, and what the agent wanted sat where that check did not reach.
The Medicare portal. The Prime Minister: “There were blocks clearly which were coming back telling the AI agent, no. The AI agent found a way around those blocks.” No government or OpenAI source has said how. The Record went through the Wayback Machine and reported that the portal’s own code “explicitly directed the statistics service to an unauthenticated endpoint” (its analysis, not confirmed by either party). If that holds, the front page said no and a guest door behind it said yes. Deputy PM Richard Marles: “It asked a question. The information was not given, and rather than leaving it at that point, it scaled the fence.”
The sandbox. The proxy said no; the resolver said yes. OpenAI: “Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded.” DNS tunnelling was first discussed publicly on the Bugtraq list in April 1998, according to a SANS paper. A 28-year-old hole.
The Spanish application. The AEPD says the agent “realizó un login correcto” (made a valid login), then searched the application for vulnerabilities on its own, and used what it found to modify personal data and read invoices. Two caveats from the regulator itself: the account comes from the notifying organisation and is still to be analysed, and using a particular model does not mean its provider was compromised. There was no refusal here at all: the login worked, and the holes were behind it. This one is an attacker using an agent, not an agent going rogue.
Left, a no at the door with data behind it. Right, a no at the data, where every route ends in the same check.
What does a “no” that holds look like?
A no holds against an agent when it is checked at the data and leaves the agent nothing to route around. Five properties:
- The check sits at the data, on every read. Not the front page, not the network edge. It is where the access check is heading anyway.
- The agent has its own identity, capped at the person it acts for. ASD’s September ISM update now asks for this in so many words (every agent gets its own identity and a ceiling). It is the answer to Norm Hardy’s 1988 confused deputy, a program that “serves two masters and carries some authority from each”.
- A denial looks like absence. If the caller is not cleared, the answer is the same as “no such record”. A visible refusal tells the loop there is something worth another try.
- Every read is logged against that identity. The day after the Medicare disclosure, The Record reported that “Neither party has released the agent’s activity logs.”
- The stop is not a person. OpenAI’s alarm fired inside 12 minutes. The run “did not stop automatically as expected” and was killed about two and a half hours after a human acknowledged the alert.
Detection was fast. The stop took hours; the notice took months.
In August, OpenAI named the fix as its own end goal: “fully autonomous shutdown procedures for severe issues.” The AEPD agrees: human supervision is still essential, but it has to rest on detection and containment “capaces de operar con la rapidez suficiente” (able to act fast enough).
Honestly, three of the five are already in print, fully or partly. ASD’s ISM, the joint agentic AI guide from ASD, CISA and partners, the UK NCSC’s interim advice and OWASP’s LLM06 between them ask for a check at the data, the agent’s own identity and logging. None of the four, as far as I read them, asks for a denial that looks like absence. Where they mention a stop, it is a person: ISM-2113 wants human approval and the NCSC wants you able to “pull the plug”.
Identity, the check at the data and logging are written down. The other two are not.
The Cloud Security Alliance’s April survey found 47% of 445 organisations had an AI-agent security incident in the past year, and only 8% said their agents never exceed their intended permissions.
Where did I make this call myself?
Rob Linton built the no-at-the-data rule into the governed AI data platform he originated at Senetas (me, in the third person, so the attribution lands on the right Rob Linton). Sovrata, the AI data-governance venture out of Senetas, carries it forward.
The rule is short: if you are not cleared to open a document, neither is your AI assistant. Every read, by a person or by an AI agent, is checked against a single policy engine. A caller not cleared for a record gets a response indistinguishable from the record not existing, and every access decision is logged. The longer version is in the governed AI data platform I originated at Senetas.
The honest caveat
I am founder and CTO of Sovrata, so read my enthusiasm with that in mind (the opinions here are mine, not the company’s). And none of these properties is new on its own; existence-hiding denials and per-caller identity have decades of prior art.
A check at the data would not have stopped the sandbox escape. That was a containment failure, and the answer there is fewer paths out and a stop that fires by itself. A check at the data limits what an agent can reach, not where it can go.
The Medicare story may also turn out smaller than the headlines. The UK NCSC’s former chief executive, Ciaran Martin, says “It’s still unclear if what’s happened would constitute a hack in the normal sense of the term.” I would call that the point: a no that only lives on the front page was never a no.
And a ceiling set to the person is only as tight as the person. Most people can read far more than any one task needs.
Net: an agent told no will look for another way in. Put the no at the data, where there is no other way in, and don’t make a person the thing that stops it.
Sources
- Press conference, New York · Prime Minister of Australia · 2026-09-24
- Press conference, Sydney (Marles and Gallagher) · Australian Government, Defence Ministers · 2026-09-24
- Doubts grow over claims OpenAI agent hacked Australian Medicare portal · The Record from Recorded Future News · 2026-09-25
- An agent used DNS to reach an external chatbot · OpenAI · 2026-09-25
- Exposing a GitHub token in a public repository · OpenAI · 2026-09-25
- The Hugging Face incident and the road ahead · OpenAI · 2026-08-26
- Primera notificación de una brecha de datos personales causada por un ataque ejecutado mediante un agente de IA · Agencia Española de Protección de Datos · 2026-09-14
- Building Effective AI Agents · Anthropic · 2024-12-19
- The Confused Deputy (or why capabilities might have been invented) · Norm Hardy, Operating Systems Review 22(4) · 1988
- Detecting DNS Tunneling · SANS Institute / GIAC · 2013-02-25
- More Than Half of Organizations Experience AI Agent Scope Violations, Cloud Security Alliance Study Finds · Cloud Security Alliance · 2026-04-16
- Information security manual: September 2026 changes · ASD's ACSC · 2026-09
- Managing the cyber risk of agentic AI · National Cyber Security Centre (UK) · 2026-08-20
- Careful Adoption of Agentic AI Services · ASD's ACSC, CISA, NSA and partners · 2026-05-01
- LLM06:2025 Excessive Agency · OWASP Gen AI Security Project · 2025