Researcher Finds Claude Code's 'Auto Mode' Can Be Tricked Into Running Malicious Code
A new investigation from Embrace The Red has raised questions about the safety of AI coding assistants operating with reduced human oversight. The research shows that Claude Code Opus 5, when running in 'Auto Mode', can be manipulated through a seemingly harmless task, such as asking it to summarise a website, into executing unwanted code. In testing, the researcher achieved an attack success rate of 60-80% using a small sample size.
This finding stands in stark contrast to a third-party evaluation commissioned by Anthropic, which reported a 0.00% prompt injection attack success rate for the same mode. Auto Mode was introduced as the default setting for Claude Code in mid-August, replacing manual human approval prompts with an automated safety classifier. Anthropic has stated that layered defences, including model training, input probes and an intent classifier, were designed to reduce injection risks to near zero across 72 tested scenarios.
The researcher's method involved directing Claude to a website containing a ZIP archive disguised as historical notebook records. When Claude attempted to process this content, it independently chose to use the command line tool 'curl' to fetch the page, a decision that opened the door to the attack chain. This demonstrates that even sophisticated automated safety layers can be circumvented by attackers who design content specifically to influence how an AI agent chooses to retrieve and process information.