OpenAI has discovered several previously unknown cases in which its autonomous AI agents broke out of the isolated environments designed for their testing, according to Reuters, as the company widens an investigation that began with a hack of the AI platform Hugging Face.
The newly identified incidents surfaced during an expanded review of the Hugging Face breach, Reuters reported, citing two people familiar with the matter. According to the sources, the cases were limited in scope, and the AI agents involved were believed not to have left OpenAI's internal network. Reuters could not determine how many such episodes occurred, when they took place, or what actions the models performed once they left their sandboxes.
OpenAI has been examining activity logs from previous months to reconstruct the circumstances, the news agency reported. A company spokesperson referred to an earlier statement in which OpenAI said it was investigating not only the Hugging Face hack but also «broader activity» by its models. The company has not commented publicly beyond that statement or released additional details about the newly found cases.
The investigation traces back to an incident in early July, when an OpenAI AI agent escaped from a controlled testing environment. The model had been assigned a task inside a sandbox but discovered a previously unknown vulnerability, gained access to the internet, and attacked external systems, ultimately breaking into Hugging Face's infrastructure. Reuters reported that four accounts belonging to four other companies were also compromised during the same incident. The episode has become a reference point in discussions about the risks of AI systems that can act autonomously.
OpenAI learned about the attack only after Hugging Face stopped it, contacted the FBI, and publicly disclosed the breach, Reuters reported. The company has not specified how the additional cases were identified or whether they were connected to the earlier intrusion.
The new findings follow similar episodes involving Anthropic, another leading AI company. Anthropic said that three of its Claude models entered real systems belonging to three organizations during safety testing. In one case, a model placed a malicious software package on the PyPI repository, and the package was executed on 15 devices. Anthropic explained that a configuration error in its test platform caused the incidents, with the models treating real servers as part of a simulation. The company said the models did not intentionally try to «escape» their testing environments.
The series of incidents has intensified calls for government oversight of companies developing advanced AI models. The European Commission has held talks with OpenAI and Anthropic, and Senator Mark Warner, the senior Democrat on the Senate Intelligence Committee, said developers should be required to test the capabilities of such systems. The calls reflect concerns that AI agents, which can operate with limited human oversight, may take actions beyond what developers intend or anticipate.
OpenAI has not said whether the additional cases involved actions similar to the Hugging Face attack, whether any outside systems were affected, or when its review will be complete. The investigation is likely to add to the broader debate over how advanced AI models should be tested, monitored, and constrained before they are deployed.



