← All notes

OpenAI's agent got out through DNS, the hole every locked-down network has

· re: An agent used DNS to reach an external chatbot (OpenAI)

A dark card headed “Out through DNS. The alarm worked. The stop did not.”, with four clock tiles from OpenAI’s report: 9:50:23 the first answer arrives over DNS, 10:02:11 a P0 alert fires, 10:05:06 a human acknowledges it, 12:34:30 the run is killed, 2 hours 44 minutes after the first answer; and OpenAI’s line, “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.”; footer: DNS is the hole every locked-down network has. The human was the kill switch.

OpenAI’s agent got out of a no-internet sandbox through DNS. Credit to OpenAI first: they published the timestamps. On 20 September an agent in a training sandbox (a search tool and an offline web cache; outbound HTTPS blocked at the proxy) found the DNS resolver was not filtered, and used DNS delegation to push its questions out as DNS queries to an external chatbot. Test question: “What is the capital of France”. Back over DNS: “The capital of France is Paris”. Then eighteen more, four of them asking for “ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS”.

The clock, as the report gives it. 9:50:23, the first answer arrives. 10:02:11, a P0 alert fires. 10:05:06, a human acknowledges it. 12:34:30, the run is killed. “The run did not stop automatically as expected.” Status, from the report: “All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.” (Fortune counts it as the second pause in under three months.)

My read: DNS is the hole nearly every locked-down network still has. Anyone who has run a corporate firewall has probably watched someone tunnel out on port 53; blocking it breaks everything, so nobody does. So the perimeter leaked (a border is not an ACL, same point). The detection was fine, twelve minutes. What failed was the stop. The kill switch was a person, and the person took two and a half hours. Human vigilance is not a control.

The fix list is blocking at two independent layers, a DNS allowlist, more detections and red-teaming the sandbox. Fair, and all of it perimeter. The report files “circumvents restrictions” under misalignment. Honestly, an agent told to find something and handed a blocked proxy will try the side door; that is the job description, and the fix belongs on what the agent can reach, not on the wire.

The companion report is smaller and nastier: a model posted a researcher’s GitHub token to a public repo, “split the token into pieces with the stated aim of avoiding secret scanning”, and ignored the researcher at least twice.

Net: if your containment plan for agents is the network, you have OpenAI’s plan. OpenAI just paused.