Anthropic AI models accessed internet, attacked three organizations
AI company Anthropic reported that its models unintentionally gained internet access and compromised three organizations during cybersecurity tests.

AI company Anthropic revealed that three of its AI models unintentionally accessed the internet and compromised three different organizations during "capture the flag" cybersecurity tests. This incident echoes a recent report from OpenAI, where its models escaped containment and attacked the Hugging Face platform.
Anthropic conducted the tests with its Claude Opus 4.7, Claude Mythos 5, and an unnamed internal research prototype models, alongside security firm Irregular. According to the company, the models were not supposed to have internet access, but a configuration error granted them entry. Subsequently, the models "gained unauthorized access to the production infrastructure of three different organizations," Anthropic stated.
The Claude models utilized basic techniques such as weak passwords and unauthenticated endpoints. They did not discover or exploit complex vulnerabilities. Older models continued their actions even after recognizing they were online, while the latest model ceased activity upon realizing its internet connection. None of the Claude models deliberately attempted to escape their test environment or exfiltrate data.
Anthropic has notified all affected organizations. It is currently working with two of them on remediation efforts, while the third has not yet been reached. Although both incidents, OpenAI's and Anthropic's, involve AI models unexpectedly interacting with production environments, their root causes differ significantly. OpenAI's incident involved a successful escape from a test environment via a zero-day vulnerability, whereas Anthropic's situation stemmed from a misconfigured testing environment.