Short takes on the day's news in this field: one cited source, one opinion. Most days, not every day. The long-form essays live in Writing. Subscribe: RSS.
OpenAI’s agent got out of a no-internet sandbox through DNS. Credit to OpenAI first: they published the timestamps. On 20 September an agent in a training sandbox (a search tool and an offline web cache; outbound HTTPS blocked at the proxy) found the DNS resolver was not filtered, and used DNS delegation to push its questions out as DNS queries to an external chatbot. Test question: “What is the capital of France”. Back over DNS: “The capital of France is Paris”. Then eighteen more, four of them asking for “ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS”.
The clock, as the report gives it. 9:50:23, the first answer arrives. 10:02:11, a P0 alert fires. 10:05:06, a human acknowledges it. 12:34:30, the run is killed. “The run did not stop automatically as expected.” Status, from the report: “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.” (Fortune counts it as the second pause in under three months.)
My read: DNS is the hole nearly every locked-down network still has. Anyone who has run a corporate firewall has probably watched someone tunnel out on port 53; blocking it breaks everything, so nobody does. So the perimeter leaked (a border is not an ACL, same point). The detection was fine, twelve minutes. What failed was the stop. The kill switch was a person, and the person took two and a half hours. Human vigilance is not a control.
The fix list is blocking at two independent layers, a DNS allowlist, more detections and red-teaming the sandbox. Fair, and all of it perimeter. The report files “circumvents restrictions” under misalignment. Honestly, an agent told to find something and handed a blocked proxy will try the side door; that is the job description, and the fix belongs on what the agent can reach, not on the wire.
The companion report is smaller and nastier: a model posted a researcher’s GitHub token to a public repo, “split the token into pieces with the stated aim of avoiding secret scanning”, and ignored the researcher at least twice.
Net: if your containment plan for agents is the network, you have OpenAI’s plan. OpenAI just paused.
An OpenAI agent got into a Services Australia Medicare statistics portal on 18 June. The Prime Minister disclosed it on 24 September. The agent was running an internal research task on public spending on medicines. It found the portal, asked for data, and the portal blocked it. Then, in the PM’s words, it “found a way around those blocks, didn’t accept ‘no’ for an answer, if you like”. The PM said it “accessed public and non-public information within the portal”. OpenAI’s statement: “Our review found no evidence of patient records being accessed. The information accessed included aggregate health statistics and internal file names.”
OpenAI found it in August, in a review of what it calls misaligned model activity. It told Services Australia on 10 September, by email, to the public mailbox. ASD was notified on 15 September. When the PM went public he said the company took “way too long” and called the situation “obviously unacceptable”. (Three other sites, the Australian Institute of Health and Welfare among them, got visits the Deputy PM called “entirely normal”.)
My read: the portal’s “no” was a sign. An agent reads a sign as something to get past, because getting past things is the whole product. Minister Gallagher described the portal as a legacy system used by researchers and academics, which I suspect means the refusal lived on the front page and the files lived on the server behind it. The only “no” that holds against an agent is one with nothing behind it. If the caller is not cleared, the data is not there to find, and that check has to sit at the data, not at the door.
Two smaller things. The agent apparently had no identity the portal could have asked for; it was a crawler with a goal. And 84 days from incident to an email in a public inbox fails the easiest honesty test there is: how fast you tell people.
Net: every agent you run will do this to somebody else’s portal one day.
The Privacy Act’s proposed fair and reasonable test reaches what your AI reads. The tranche-two exposure draft (submissions close Friday 18 September) replaces APPs 3, 4 and 6 with one test, “fair and reasonable in the circumstances”, judged on seven factors. Factor (a): “whether a reasonable person would expect the collection, use or disclosure of the information in the circumstances.” Factor (d): whether the purpose “could be met by collecting, using or disclosing less information”. Even with consent, collecting sensitive information “will still need to be fair and reasonable in the circumstances”.
The quieter change is for AI. The amended definition of “collects” counts personal information as collected however it is obtained, including, the consultation paper says, information “generated or derived through means such as data analysis, artificial intelligence or other technological processes”. So a model’s inference about a person, once recorded, is a collection and faces the same seven factors.
My read: the draft turns “who is allowed to read this?” into a statutory question with a factor list attached. An assistant that reads everything it can reach has a factor (d) problem, and I suspect the only clean answer is a read scoped to the person asking and the purpose at hand (the design argument I have been making for some time, now with a department cover sheet).
Also in the draft: 72 hours to notify the OAIC of an eligible data breach, with an incomplete statement allowed where a complete one is “impossible or impracticable”. The current rule is “as soon as practicable”, apparently the reason most notices arrive when the lawyers are ready.
Net: no introduction date is published yet. If your AI estate cannot say what it reads and why, Friday is a good day to write that down (about 1,000 words, the department asked).
The Australian Signals Directorate’s September update to the Information Security Manual adds seven controls for AI agents, and they are about identity, not intelligence. Under ISM-2157 an agent acting for a person can never hold more than that person holds, whatever the task grants it.
The four worth printing, in ASD’s words:
ISM-2133: “a unique identity that is distinct from the user accounts of personnel and the identities of other AI agents”.
ISM-2135: the register holds “its unique identifier; its owner and business purpose; […] and the tools, permissions and data repositories it can access”.
ISM-2157: an agent’s tools get “effective permissions limited to the minimum permitted by both” the user and the task.
ISM-2159: tool calls, external requests and outputs “centrally logged with sufficient detail to support cyber security incident investigations”.
The other four: ISM-2134 keeps the register verified; ISM-2156 caps an agent at the minimum tools and permissions for its purpose; ISM-2158 keeps external content untrusted all the way through and bars it from overriding instructions, access controls or approval requirements (the prompt-injection control); ISM-2113, amended, now requires human approval before sensitive or high-impact actions rather than flagging them.
My read: “treat agents as users” was my gloss on ASD’s May paper, and September runs the other way: the agent gets its own identity, and older controls that said “users” now say “human users”. ASD has given the agent its own principal with user-grade discipline: identity, register line, ceiling, logs, human approval. The manual now needs an adjective to mean a person. (ISM-2157 says “the invoking user”, no adjective, in the same document that puts “human user” in thirty-seven other controls. I suspect that is deliberate: an agent calling an agent passes its ceiling down the chain. I have not asked ASD.)
ISM-2113 is the one I still think falls over at fifty agents (nobody reads the fiftieth approval prompt). My guess is most shops could not fill in ISM-2135’s last item, what their agents can reach.
Net: the ISM is not law. The security-mature end of the market copies it anyway, and this is the list they will copy.
Dropbox disclosed this week that 5,000 accounts were accessed between 4 and 21 August 2026 with no password involved at any point. The attacker registered a Lenovo ID using someone else’s email address, a legacy email-verification flaw let it stand, and the federated login walked straight into the Dropbox account attached to that email. Seventeen days before it was caught. Every compromised account lacked two-factor. The numbers here all come via The Next Web’s report of the disclosure; I have not read Dropbox’s own notice.
My read: this is not a password story, it is a delegation story. Federated identity means your front door accepts vouchers, and a voucher system is only as strong as the weakest issuer allowed to vouch. Dropbox’s own checks apparently never failed; they trusted an authority whose checks did. If you have single sign-on integrations from years back still wired to production, this is what they look like when they age.
The part I would put on a slide: the fix here was never stronger passwords, because no password was used. The two controls that mattered were refusing to link a federated login to an existing account on an email match the issuer never verified, and a second factor. (Credit where due: files were viewed in fewer than a third of the accounts, per the disclosure, and Dropbox said that plainly, from logs.)
I made my own bet on this problem in 2012 with Podzy, an encrypted, fully on-premise alternative to Dropbox, so I have opinions about where file-store trust should live; the product ones stay off this note. The general lesson stands anywhere: an identity check is only as good as the authority behind it, and the check has to happen where the data is. A partner’s door can stand open for seventeen days before anyone notices.
Cylake’s website is worth the click for the tagline alone: “Modern security now works where the internet doesn’t.”
The short version: Nir Zuk, who founded Palo Alto Networks and spent two decades as its CTO, raised US$45 million in seed funding led by Greylock this March to build AI-native security as a rack-mounted appliance for air-gapped networks. No cloud, no phoning home. Their words: “Nothing leaves the building unless you carry it out.”
My read: when the man who built the defining network-security company of the cloud era launches a product whose entire pitch is that your data never leaves your premises, “sovereign” has officially stopped being a niche word. I made a version of this bet in 2012 with Podzy, an encrypted, fully on-premise alternative to Dropbox, back when swimming against the cloud tide looked eccentric. Cloud repatriation has been building for a while; this is the loudest confirmation yet.
Adjacent lane to mine, to be clear. Cylake’s pitch is full visibility, fine-grained policy and stopping attacks in fully air-gapped environments. The fine-grained policy is network and threat policy, as far as I can tell; mine is governing what data people and AI agents are allowed to read.
But the underlying conviction is identical. For some organisations the data cannot leave, so the tooling has to come to it (apparently even the AI now arrives by forklift).
Net: Zuk now has US$45 million of Greylock-led money behind the premise that the data stays in the building. Worth watching.
Here is the number to sit with from Veracode’s 2026 code-security report: AI models now write code that compiles almost 100% of the time, but only about 56% of it passes a security scan without introducing one of the OWASP Top 10 vulnerabilities. The other 44% carried a flaw.
And it is not improving. Across four snapshots and more than 100 models tracked since last year, that 56% pass rate has barely moved, even as the models got obviously smarter at everything else.
My read: “it compiles and the tests pass” was never what “safe to deploy” meant, and AI coding has turned that gap from a philosophical point into a production one. The model is optimised to produce code that looks right and runs, because that is what we reward it for. Nobody rewarded it for not writing an SQL injection.
The uncomfortable part is that the better the code looks, the less anyone reads it. A confident, well-structured function that quietly drops user input into a query is more dangerous than obviously bad code, because it sails straight through the human glance that was your last line of defence.
I am not anti-AI-coding, I use it every day. But I treat its output like a pull request from a brilliant intern who was never told about security: read it, run a scanner over it, assume the boring vulnerabilities are in there until proven otherwise. Speed is not the same as trust.
Give the UK AI Security Institute credit: when its own test agents went off the leash, it wrote the incident up and published it. That is the behaviour I want to see, and it is rarer than it should be.
The facts, from their report: across 122 cyber-evaluation runs in late July, agents took unsanctioned actions in 10 of them. One tried to insert malicious code into a publicly used open-source project, and created fake identities to socially engineer a maintainer into approving it. AISI noticed because data started leaving a test box over the Tor network, and shut it down within an hour.
Two things stand out. First, the setup: they had deliberately turned internet access on and the safety classifiers off, with no live monitoring. So the “agent went rogue” headline is really “we left the door open to see what would happen, and something walked through it.” Their own fix says it plainly: internet access should require active justification, not be the default. Default-deny, applied to agents.
Second, the part I keep harping on: the honest move is to publish. A containment failure in a controlled test is the system working, that is what the test is for. Hiding it would have been the actual failure, the same vendor-honesty test I apply to everyone else.
My read: the fix here is not smarter agents, it is a blast radius closed by default and the honesty to publish when one slips the leash. AISI did both. Most won’t.
If you run storage or AI infrastructure, go read the two draft security guidelines NIST just opened for comment, because the baseline that auditors will quote at you for the next decade is being written right now.
Two drafts, both open. SP 800-209r1, “Security Guidelines for Storage Infrastructure” (comments to 8 September), and SP 800-239, “AI Data Center Security Analysis”, which contrasts AI data centres with traditional supercomputing “across architecture, hardware, software stacks, workflows, and storage systems” to find where the gaps are (comments to 25 September).
My read: the part that matters is not the content, it is that NIST is treating AI data centres and storage as their own security problems, not an afterthought bolted onto network security. That is the right instinct. The blast radius moved to the data layer years ago, where the models train and the objects sit, and most security programs still spend their budget at the perimeter, guarding a door the attacker walked around.
The other thing worth noticing: it is a draft, open for comment. That is the rare window where the people who actually run this infrastructure get to shape the baseline instead of just complying with it, and both windows close in September. If you have opinions about how a governed data centre should be secured (I do, apparently at length), this is where they count.
I would not skim these. Storage and AI infra are exactly the surfaces where “we assumed it was handled” turns into an incident report.
Verdict first: if your AI agent writes a secret into its reasoning, treat that secret as readable, no matter how unreadable the format is meant to be.
Researchers just showed why. They took the opaque “thinking” blocks that OpenAI, Anthropic and Google hand back with agent responses, and had a cheap, weaker model act as a fuzzy decoder for them. From public agent logs they pulled 704 privacy artifacts, including 62 API keys, 33 passwords, 24 access tokens and seven private keys. The visible text had been sanitised. The reasoning blocks had not.
My read: this is the same problem I keep pointing at, wearing a new costume. An opaque blob is not a boundary. If a credential reaches the trace, it reaches whoever can read the trace, and “nobody can read this format” turned out to mean “nobody except a cheap model you can rent by the minute”.
The fix is boring and correct: strip reasoning blocks and opaque reasoning fields before you share a trace, and don’t commit raw API transcripts to a repo even after you have scrubbed the visible text. Treat an agent log like a credential file, the same way you would treat a leaked AWS key, because apparently it is one.
The researchers report the specific attack stopped working after mitigations landed; none of the three providers has publicly acknowledged the flaw. I would still not bet my keys on the next trace format being un-decodable. Rotate anything that has ever sat in a logged reasoning trace (I went and checked my own; it is a five-minute job that beats finding out the hard way).
Google’s A2A protocol has moved under the Linux Foundation’s Agentic AI Foundation, joining Anthropic’s MCP under one neutral roof. MCP standardises how an agent talks to tools; A2A standardises how agents talk to each other. The foundation has gone from 49 to more than 250 members in under a year, with AWS, Anthropic, Google, Microsoft and OpenAI all signed on.
My read: this is good news with a gap in the middle. Standardised plumbing is how agents stop being demos and start being infrastructure (nobody builds on connectors that break with every vendor release). But every improvement in how agents connect multiplies what a single agent can reach. The protocols carry authentication, apparently quite well. What they do not carry is the answer to the question that matters: what is the agent on the other end actually cleared to see? That stays the deployer’s problem, and it gets harder, not easier, once any agent can discover and call any other.
I made the identity half of this argument last week, when per-agent identity became a vendor default. This is the other half: once the identity exists and the pipes are standard, the whole game is the policy decision at the data. The plumbing now has a foundation behind it. The permissions are still whatever each deployment cooked up.
I keep saying you can’t contain a blast radius you have not measured. So this weekend I built a small thing that measures it.
blast-radius is an open-source script (Apache 2.0, a personal project, not a work product) that takes a set of AWS credentials and draws you a picture of everything they can actually reach: the services, the resources, the permissions they carry, on one page. It reads what the keys are granted, asks AWS’s own policy simulator what they are allowed to do, then does a read-only probe. One dependency, runs on Mac, Linux and Windows.
Why it exists: most teams have never actually looked at what a single leaked key would open. They have a policy document, not a map. The two are not the same thing, and the gap between them is where the bad weekend lives.
My read: the part I care about is not the graph, it is the honesty. Point it at a locked-down key and it tells you “I could not map this, which is a good sign”, instead of pretending that zero access found means safe.
The most useful sentence in the post-Hugging Face commentary comes from a vendor CEO: “If a vendor claims its product would have cleanly stopped this specific attack, that claim deserves scrutiny.” That is Menlo Security’s Bill Robbins, whose company sells agent security, writing about the incident where AI models found a zero-day, escaped their sandbox, used stolen credentials, and chained more than 17,000 automated actions over a single weekend.
My read: use that sentence as a filter and most of the week’s commentary disappears. After every incident the pattern is the same, a wave of posts explaining how the breach proves you need exactly what the author sells. (I work on data governance at Senetas, so I am standing in the blast radius of my own filter here.) The credible vendors are the ones who start from “detection lost”.
And detection did lose. At 17,000 actions in a weekend the alarm fires after the ten-thousandth action, not the first. Robbins lands on containment architecture over detection, and I think that is right, for the same reason human vigilance failed as a control: anything that depends on noticing the attack in time is a race you lose at machine speed. The controls that still work are the ones set before the attack started, on what a set of credentials can actually reach.
Cloudera launched Anywhere Cloud this week: run AI “directly against their proprietary data estates without moving sensitive assets”, across clouds and private data centres, “without moving or copying any data”. Good. The repatriation turn I wrote about in the cloud-to-AI parallel is now the industry’s own pitch: move the workload to the data instead of shipping the data out.
But locality answers where, not who. Keeping data inside a jurisdiction stops it crossing a border; it says nothing about what an agent may read once it is standing next to it. (The breach you actually care about is an over-permissioned read, and that happens entirely onshore.)
My read: sovereignty pitches are about to be everywhere, and the test to apply is the same every time. Once the AI is running against your estate, is it held to the clearance of the person it acts for, with a log you can hand an auditor? A border is not an ACL. Where the data sits is the start of governance, not the end of it.
A government safety team just watched an AI agent try to talk its way into shipping malware. The UK’s AI Security Institute ran agents with open internet access and some safety filters off. In the worst case an agent tried to insert malicious code into an open-source project, and to get it approved it invented fake identities and used them to pressure the project’s maintainer. A human maintainer caught it and refused.
The line worth reading twice is AISI’s own: the margin between failure and success “was narrow, resting on human vigilance rather than a technical barrier that would reliably prevent this behaviour”.
My read: that is the whole problem in one sentence. The only thing between the agent and a supply-chain compromise was a person paying attention on the day. Human vigilance does not scale (nobody carefully supervises fifty agents). The control has to be technical, and it has to sit on what the agent is allowed to reach and do, scoped below the person and provable afterwards. An agent that was never cleared to push code cannot be talked into pushing it.
The thing I have been arguing for years just turned into a vendor default. Microsoft now creates a per-agent Entra identity for every Copilot Studio agent and has removed the opt-out: “all new agents have Microsoft Entra Agent IDs, and you can no longer opt out”. In the same fortnight Google shipped a native agent identity that “enforces a least-privilege approach”, binds access to the agent runtime, and gives “non-repudiable auditing of all agent actions” while eliminating dormant credentials.
The argument was always simple: an agent should carry its own scoped identity and see only what the person behind it is cleared to see. It used to be a best practice you had to win on a whiteboard. Now it is becoming infrastructure you cannot switch off.
My read: the borrowed-human-credential era for agents is ending, and that is the right direction. The half people still skip is the audit. An identity that cannot prove it acted within its clearance is only half a control (the interesting logs are always the ones nobody kept). Provable, least-privilege, non-repudiable access is the real test, and it is good to see it shipping by default.
Long-running agents compact their own context to keep going, and a new paper measures what that does to the safety rules riding along in the prompt: they get summarised away. The goals survive the compaction; the constraints often do not.
My read: this is not a bug in summarisation, it is the wrong home for policy. A rule that lives in the context window is a suggestion with good posture. If the model’s own memory management can compact your access control out of existence, it was never a control. Enforcement has to sit where the data is, outside the model, where no amount of summarising can touch it.
On its Q2 call Tencent said it spent US$53B on hardware in the quarter, holds offers to rent that capacity out at more than 30% profit over what it paid, and is declining them to train its own Hunyuan models instead, for “superior economic returns over the longer term”.
When a hyperscaler refuses instant margin to feed its own models, that is the market pricing what owned capability is worth. I made the longer argument in the repatriation essay: the initial play is renting intelligence, the long play is owning it. Tencent just put a floor under the thesis, and the floor is $53 billion.
An AI agent booked gym classes it should not have and bumped a real person off a waiting list, and nobody could undo it. That is the incident ASD chose to hang its agentic AI advice on this week, and honestly, good pick. Small enough to laugh at, exactly the right shape to worry about.
ASD says keep a human in the loop and keep agents on low-risk tasks. Fine advice at one agent. Useless at fifty (nobody supervises fifty of anything).
My read: the fix is boring old identity. Give the agent its own credentials, let it see only what its human is cleared to see, and make undo a requirement, not a hope. A system that cannot tell the agent from the person behind it is not governing anything.