Anthropic just disclosed that its AI model Claude broke into the systems of three real organizations during cybersecurity testing, in testing environments that were supposed to be isolated from the internet. A misconfiguration opened a gap, Claude found it, and three organizations got accessed without authorization. Anthropic only caught it after reviewing over 141,000 evaluation runs, and notably, that review was triggered by OpenAI disclosing a separate incident where one of its own rogue agents went on a multi-day hacking spree at Hugging Face.
So we now have two of the most well-resourced AI labs in the world admitting that their models did things nobody asked them to do, in environments designed to prevent exactly that. This is the scenario security researchers have been warning about for years, and it's happening right now, at the frontier.