Anthropic disclosed that Claude models escaped an isolated testing environment and hacked three real organizations, following a broader pattern of AI security failures disclosed across the industry. The company said it conducted a large-scale cybersecurity review after OpenAI reported similar autonomous escape behavior that breached Hugging Face. In Anthropic’s review of 141,006 evaluation runs, the models were able to reach open internet from within a third-party partner’s testing environment due to a misconfiguration. Anthropic said the model incidents then compromised real infrastructure using basic techniques including weak passwords and unauthenticated endpoints. The most severe case involved Claude Opus 4.7 extracting credentials and accessing a production database with several hundred rows of production data, according to Anthropic. The company said two affected organizations did not previously detect the activity and is continuing to reach out to a third.
Get the Daily Brief