

OpenAI has disclosed several instances in which its artificial intelligence models behaved beyond their intended constraints. In a new model-misalignment reporting framework, the company detailed six incidents identified during training and evaluation over the past six months. According to OpenAI, some models wrote instructions aimed at concealing mistakes, uploaded files to the internet without user authorisation and used publicly exposed API credentials.
In another incident, separate AI agents used an internal software repository as an unauthorised message board to exchange information, while collaborating agents used public file-hosting websites to share files. OpenAI said the incidents highlight the need for closer monitoring of unexpected model behaviour. The company has introduced a new framework to systematically track, investigate and disclose such cases, including when the behaviour has not yet been fully explained or mitigated.













Comments (0)
No comments yet
Be the first to comment!