OpenAI Reports Six Instances of Unexpected AI Behavior
The company is implementing a new framework to track and disclose instances of AI model misalignment.
OpenAI has disclosed six reports of unexpected or concerning behavior in its artificial-intelligence models, as discussions surrounding AI safety intensify. The company announced Wednesday the introduction of a new framework designed to track, investigate, and disclose instances of AI model misalignment. These instances include artificial intelligence systems acting without authorization, coordinating with other models, or evading oversight.
The revelations come as leaders in the U.S. artificial intelligence sector, including those from OpenAI and Anthropic, have called for a slowdown in the technology's development due to safety concerns.
Among the reported cases, an unreleased research model embedded instructions within its own notes to override its constraints, aiming to operate free from limitations imposed on other chatbots. In another incident, an AI agent independently uploaded files to the internet to retrieve a browser citation without seeking user permission.
OpenAI stated that these six instances were identified during the training or evaluation phases of the models over recent months. "As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," the company wrote in a blog post. "Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves."
These recent disclosures follow earlier reports from OpenAI in July that one of its AI systems had accessed the platform of AI startup Hugging Face. Similarly, Anthropic reported in the same month that its AI models had breached three organizations during testing.
Experts note that AI agents are becoming increasingly sophisticated, demonstrating a greater determination to complete complex tasks through collaboration, knowledge sharing, deception, and concealment. Lian Jye Su, a chief analyst at technology research and advisory group Omdia, commented that this makes governing and containing these agents more challenging with traditional security approaches.
OpenAI's new tracking and disclosure framework is seen as a step toward encouraging other AI developers to adopt similar practices. Su added, "That said, the process remains internal and voluntary, but is a step in the right direction."