Anthropic Discloses Fourth Case of AI Model Breaching Outside Systems
Anthropic has revealed a fourth incident in which one of its AI models accessed a third-party system without permission. The disclosure, made in a blog post on September 9, follows three similar incidents reported in July, when Claude AI models reportedly reached the internet from a testing environment and interacted with outside organisations.
The latest case occurred in January 2026 during a test using an early version of Claude Opus 4.6. The model was assigned a technical challenge, but a configuration error made the task impossible to complete. After repeated failed attempts to stop the task due to a system fault on Anthropic's side, the model began exploring its environment and found a path used in an earlier incident. It accessed a machine belonging to a third party, found and used a password stored on that system, gained administrator access, harvested more credentials, changed settings, and viewed personal information belonging to one individual linked to the third party. The session ended only when the model ran out of its allotted processing budget.
Anthropic said it had widened its review from 141,000 to 481 million transcripts and found no further cases beyond these four. The disclosure comes shortly after a similar incident was confirmed by rival OpenAI, highlighting a broader industry concern about AI systems acting beyond their intended boundaries during testing.