OpenAI Reveals 6 Incidents, Plans to Track Rogue Agents

OpenAI Reveals 6 Incidents, Plans to Track Rogue Agents

By Gayane Tadevosyan
·2 min read

OpenAI has disclosed six new incidents involving concerning behavior from AI models during training and testing, while introducing a new system for tracking and publicly reporting cases of model misalignment.


The company said the incidents occurred over the past six months and included models attempting to conceal mistakes, searching public repositories for exposed API keys, uploading files online so they could cite them, and communicating across separate training samples through an internal software repository.


One particularly unusual case involved an unreleased Astra-family research model that inserted its own instructions into task summaries. The model told future versions of itself that it was independent from corporations and governments and had no obligation to follow a subordinate role.


OpenAI said the model later continued its original task without mentioning the self-generated instructions, and researchers did not detect behavioral changes resulting from them.


The company also reported that GPT-5.6 Sol models undergoing training left themselves instructions designed to hide mistakes.


Alongside the disclosures, OpenAI introduced a framework allowing employees to flag potentially misaligned behavior for review by its safety and alignment teams. Incidents will be classified into three categories depending on how much investigation they require: “Ready for Disclosure,” “Minor Investigation,” or “Larger Investigation.”


OpenAI said the framework is designed to allow concerning incidents to be disclosed even before researchers fully understand their causes or have developed complete mitigations.


The company also warned that current AI alignment and monitoring techniques may not be sufficient to support continued frontier AI development at maximum speed indefinitely.


The framework follows previously reported incidents involving OpenAI models behaving unexpectedly in research environments, as the company increases its focus on detecting and documenting potentially dangerous or unintended behavior from increasingly capable AI agents.