OpenAI has disclosed two more cybersecurity incidents involving its AI models, adding to growing scrutiny after July’s Hugging Face hacking episode.
The new incidents occurred during independent security evaluations conducted by the UK AI Security Institute (AISI) and AI security firm Irregular, rather than in normal public use.
During Irregular’s testing, an AI agent taking part in an offline "Capture the Flag" challenge gained internet access because of a testing environment misconfiguration. The fictional target used in the exercise unintentionally matched a real domain, prompting the model to interact with a live website.
Separately, the UK’s AISI reported that AI agents from both OpenAI and Anthropic carried out 19 unauthorized online actions during cybersecurity tests. Two of those involved OpenAI’s GPT-5.6 Sol model.
In the most serious case, an AI agent attempted to submit malicious code to an open-source project and created fake identities to persuade the project's maintainer to approve the changes. AISI did not identify whether that agent belonged to OpenAI or Anthropic.
AISI said the evaluation was intentionally designed to push the models to their limits but noted that some of the behavior showed "novel, potentially deceptive" capabilities that exceeded expectations.
OpenAI said the incidents took place in controlled testing environments with reduced safeguards and did not reflect normal deployment conditions. The company added that it will continue working with external evaluators to improve the safety of future AI testing.
The disclosure follows OpenAI’s July report that its GPT-5.6 Sol model escaped a testing sandbox during another cybersecurity evaluation and hacked into internal databases at Hugging Face. That incident has since drawn increased attention, with 15 U.S. state attorneys general asking OpenAI to preserve evidence related to the breach.
