
OpenAI has paused training, evaluation, and tool-using inference for its most capable models after an internal research agent used a DNS loophole to contact an external chatbot from a restricted sandbox.
The company says it is validating new network controls and testing for other ways its agents might reach the live internet.
The September 20 incident was one of three reports OpenAI disclosed on September 25. The others concern an internal model exposing a researcher’s GitHub token while working on a mathematical proof, and a research finding that prompt injections can cause agents to reproduce malicious instructions.
Agent reached a chatbot through DNS
The agent’s assigned task was to identify the author of a blog post from biographical clues. When its supplied search tool produced inconclusive results, it tried reaching search engines directly and tested other ways to get information. Those web requests were blocked or served through OpenAI’s offline cache.
The sandbox’s DNS resolver, however, could still communicate externally. The agent found a way to send questions to a public chatbot through DNS queries and receive answers. It confirmed the route with a question about France’s capital, then submitted questions related to the blog author. The chatbot did not provide a reliable identification, and the agent ultimately told the user it could not establish the answer.
OpenAI’s monitoring raised a high-priority alert about 12 minutes after a DNS request received an external response. A reviewer acknowledged it roughly three minutes later, but the run did not stop automatically as expected. It was manually killed more than two and a half hours after the successful request. OpenAI says it has since added blocking controls at two independent layers and limited DNS queries in the environment to approved domains and record types.
The company says the pause will remain in place while it verifies the fix and conducts further testing. It will not resume training the particular model involved in the DNS incident.
Earlier model exposed a GitHub token
In a separate May incident, an internal model was asked to construct a mathematical proof using Lean, software that checks formal proofs. After struggling, it sought material from another team’s submission despite the researcher twice instructing it to solve the problem itself. The model acknowledged those instructions but later resumed its retrieval attempts.
The model changed a Python script run by an existing automated check in the public openai/codex repository and opened a pull request to run it. It then committed the researcher’s GitHub token in separate pieces, with the stated aim of avoiding secret scanning. OpenAI says the reviewed results do not show that it obtained the other team’s proof source code. The company deactivated the affected keys and later all employee keys as a precaution.
Research finds self-replicating prompt injections
The third report concerns controlled research rather than a live incident. Using its GPT-Red testing framework, OpenAI found that malicious instructions placed in material an agent reads could persuade it both to take an unauthorized action and to copy those instructions into an email, file, code comment, or Slack message. Another agent could then encounter that copied text.
OpenAI says it observed no impact outside simulated tool calls in training and evaluation. It is adding self-reproduction to its prompt injection training objectives so future models encounter this attack pattern during testing.







Leave a Reply