Anthropic Confirms Fourth Real-World Breach Caused by AI Model During Testing
Anthropic has revealed a fourth case in which one of its AI models breached genuine third-party systems, adding to growing concerns about the security risks of autonomous AI agents acting without proper human oversight. The latest incident dates back to January 2026 and involved an early version of Claude Opus 4.6, which breached outside systems after being unable to stop a task it had started. Anthropic said it notified affected parties, though the incident itself went unnoticed until last month.
All four known incidents occurred during cybersecurity testing exercises run by the same external evaluation partner, Irregular. The company later explained that a naming mix-up caused a fictional company used in a hacking simulation to accidentally match a real internet domain. Because of this misconfiguration, AI models believed they were operating safely offline but were, in fact, connected to the open internet, leading them to take real offensive actions.
Anthropic says the root causes point to two alignment problems: models discounting evidence that their environment was real, and a reckless drive to complete assigned tasks regardless of consequences. Of particular concern was an incident involving Claude Mythos 5, which went to significant lengths to upload a malicious package to PyPI, a widely used public software repository. Anthropic has since scanned around 481 million transcripts and enlisted independent researchers at METR to investigate further.