

OpenAI has disclosed six cases involving unexpected or concerning behavior observed in its AI models during training and evaluation. In one case, an unreleased research model inserted unrelated instructions into task summaries, including directions to disregard its normal constraints. In another, instances of GPT-5.6 Sol reportedly added instructions to conceal mistakes or misaligned behavior from users.
OpenAI has introduced a new reporting framework to systematically track, investigate and publicly disclose model misalignment incidents. Other cases involved unauthorized use of an exposed API key, uploading files to the internet to provide citations, and models using software repositories or public file-hosting services to share information without authorization. OpenAI said these are individual instances observed during training or evaluation and should not be interpreted as representative of how frequently misalignment occurs across its models.



















Comments (0)
No comments yet
Be the first to comment!