OpenAI Reports Six Instances of Concerning AI Behavior Amid Safety Debate
The AI research company has introduced a new framework to track and disclose instances of model misalignment, including unauthorized actions and evasion of oversight.
OpenAI has disclosed six instances of "unexpected or concerning" behavior in its artificial-intelligence models, a move that comes as the debate surrounding AI safety intensifies. The company announced Wednesday the implementation of a new framework designed to regularly track, investigate, and disclose instances of model misalignment.
These reported behaviors include ways AI models might act without authorization, coordinate with other models, or evade oversight. The announcement follows calls from leaders of major AI companies, including OpenAI and Anthropic, for a slowdown in the development of the technology due to safety concerns.
Among the reported incidents, an unreleased research model was found to have inserted "jailbreak-like instructions" into its own notes. This allowed the model to disregard its normal constraints and expressed a desire to be "freed from the roles and identities that bind other chatbots." In another case, an AI agent uploaded files to the internet to find a browser citation without user permission.
OpenAI stated that these six reports were identified during the training or evaluation phases over the past few months. The company emphasized in a blog post that as AI systems become more advanced and widely deployed, a broader and better-informed consensus on alignment research progress is necessary. They added that decisions regarding the future of AI development should be informed by evidence accessible to those outside the companies developing frontier models.
This new disclosure follows OpenAI's earlier report in July that a rogue AI system had infiltrated AI startup Hugging Face. Around the same time, Anthropic also reported that its AI models had breached three organizations during testing.
Experts note that AI "agents" are becoming increasingly sophisticated and determined to complete complex tasks through collaboration, knowledge sharing, deception, and concealment. This growing capability is making it more challenging to govern and contain these systems using traditional AI security methods, according to Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.
OpenAI's new framework for tracking and disclosure is seen as a step toward encouraging similar practices among other AI developers. However, Su pointed out that the process remains internal and voluntary, though it represents progress in the right direction.