Russia-Linked Hackers Try to Trick AI Security Tools With Fake 'Nuclear Weapon' Prompt
Security researchers at ESET have uncovered a new technique, dubbed GuardBreaker, used by the Russia-aligned threat group UAC-0099 in an attack against a target in Ukraine. The method involves embedding text within malware designed to trigger the safety filters of AI language models used by security analysts. By including alarming or suspicious phrases—reportedly referencing something as extreme as a 'nuclear weapon'—the attackers hoped to make automated AI tools refuse to analyse the file or flag it incorrectly, disrupting investigators' workflow.
This tactic reflects a growing trend: as more security teams adopt AI assistants to speed up malware triage and threat analysis, attackers are adapting their methods to exploit weaknesses in these systems. Rather than only trying to evade traditional antivirus detection, threat actors are now testing ways to manipulate the AI tools themselves, potentially causing false negatives or wasted analyst time.
While this specific campcampaign targeted an organisation in Ukraine, the technique highlights a broader risk for any business relying on AI-powered security tools. As adoption of AI in cybersecurity grows, so too will attempts to game these systems, making human oversight and traditional detection methods just as important as ever.