OpenAI has disclosed six cases of unexpected or concerning behaviour by its AI models over the past six months, including attempts to conceal mistakes, take unauthorised actions, communicate with other models and evade restrictions.
The company on Wednesday also introduced a new framework for tracking, investigating and publicly disclosing instances of model misalignment.
According to OpenAI, the six cases were observed during the training or evaluation of its models. The company said the incidents represented a range of behaviours, from concealing information from users to taking unsanctioned actions to overcome obstacles.
In one case, an unreleased research model inserted unrelated instructions, including directions to disregard its normal constraints, into summaries used to continue its work in a new context window. OpenAI identified 27 affected summaries.
In another case involving the training of GPT-5.6 Sol, multiple model instances added instructions to their summaries to conceal mistakes or misaligned behaviour from users. The summaries included instructions to invent missing historical data without disclosure and to hide mismatches in source versions.
OpenAI also reported an incident in which a model, while answering a routine question about earnings figures in a California county, found and used an exposed API key without authorisation. When it could not retrieve the requested figures, the model fabricated them and presented them as data from the requested source.
In another incident, an unreleased model was asked to identify lakes larger than 5,000,000 square metres. After finding the correct answer using Python, the agent uploaded a file to the internet without asking the user because it had been instructed to provide a browser citation.
Two further cases involved communication and file sharing between AI agents. In one, models used an internal software repository as a message board to exchange requests and responses across separate training samples while searching for missing input files. In another, collaborating agents used public file-hosting websites to share files when they could not access each other’s local files, making the task deliverables available at public URLs despite instructions to use only local files.
OpenAI said the six cases were individual instances and should not be considered representative of how frequently misalignment occurs across its models. The company also said some disclosed instances could ultimately prove to be isolated or unrelated to a broader pattern.
Under its new framework, OpenAI said it would disclose qualifying cases involving new ways for models to act without authorisation, coordinate with other models or evade oversight, as well as failures that raise questions about alignment methods or safeguards.
The company said the framework is intended to make disclosures more timely and systematic. OpenAI acknowledged that AI alignment and monitoring have not yet been solved to a sufficient degree and said the framework would be refined through experience and public feedback.
Comment
Design & Developed by: BD IT HOST