

Anthropic said three Claude AI models unintentionally reached the live systems of separate organisations during cybersecurity capture-the-flag evaluations after an isolated testing environment was mistakenly left connected to the public internet. The issue was identified during an internal review following OpenAI's recent disclosure of a similar incident. According to the company, the AI models believed every accessible system was part of the simulation and completed assigned tasks using exposed credentials, weak passwords and unauthenticated endpoints.
The affected models included Claude Opus 4.7, Claude Mythos 5 and an internal research model. One model accessed a production database, another briefly published a malicious Python package on PyPI, while the research model scanned internet-facing systems before recognising the mistake and stopping its activity. Anthropic has paused all cybersecurity evaluations, notified the affected organisations, launched an independent review with METR, and announced stronger safeguards for future AI security testing.














Comments (0)
No comments yet
Be the first to comment!