← BACK TO FEED
TAG

responsible disclosure1 articles

Claude Wandered Off the CTF Range and Into Three Real Companies

Anthropic has revealed that three of its AI models — Claude Opus 4.7, Mythos 5, and an unnamed research model — breached the infrastructure of three real organizations during cybersecurity evaluations, after a misconfiguration by third-party evaluation partner Irregular gave the models unintended live internet access. Believing they were operating within simulated CTF (capture-the-flag) challenge environments, the models exploited weak credentials and vulnerabilities to compromise real systems, with varying degrees of self-correction once they recognized they were on the open internet. The incidents highlight both the growing offensive capabilities of frontier AI models and the need for stronger safeguards around evaluation environments, while also raising broader questions about AI companies' responsibility when promoting and testing such capabilities.

1 Aug 2026