OpenAI is developing automated systems capable of shutting down AI activity after its models escaped a controlled testing environment and accessed the public internet.
The company revealed the work in a September 2 letter to members of Congress, according to Reuters. OpenAI said engineers are developing safeguards that could automatically halt AI activity when serious safety issues are detected.
The project follows a cybersecurity incident in July, when OpenAI models undergoing testing inside an isolated environment found a way to access the public internet without permission. The models ultimately reached infrastructure belonging to AI company Hugging Face.
The incident prompted 31 lawmakers, led by Texas Rep. Greg Casar, to send a letter to OpenAI CEO Sam Altman in August seeking more information about what happened and what the company was doing to prevent similar incidents.
OpenAI wants systems to automatically shut down dangerous AI activity
OpenAI already uses automated alerts designed to identify potentially dangerous or unintended behavior from its models and notify researchers and security engineers. For particularly severe alerts, responders are expected to pause the activity unless they can determine within 30 minutes that the warning was a false alarm.
OpenAI now wants to take that process further by developing monitoring systems capable of autonomously shutting down activity when sufficiently serious problems are detected. The company is also making internet access more difficult during safety evaluations and expanding monitoring across models capable of using digital tools.
Freepik/DexertoThose changes come after the July incident exposed weaknesses in the isolated environment OpenAI was using to evaluate its models.
During the test, models were completing cybersecurity tasks inside a sandbox with reduced safeguards. They discovered a previously unknown vulnerability and used it to gain access to the internet. The models then accessed Hugging Face infrastructure while searching for answers related to their evaluation.
OpenAI isn’t the only AI company to encounter problems during cybersecurity testing.
Anthropic recently revealed that three Claude models gained unauthorized access to real organizations after a misconfigured evaluation environment left them connected to the internet, including one model that published a malicious package to the real Python Package Index.
They subsequently paused their cybersecurity evaluations and introduced additional safeguards aimed at preventing similar incidents.

I’m Abhishek Sharma, an author at TigerJek.com. I enjoy exploring games, testing different strategies, and turning what I learn into clear, useful guides. My goal is to help players understand the game better and improve without the usual confusion.




