New Attack Technique Tricks AI Chatbots Into Ignoring Safety Rules
Security researchers have identified a new technique that can bypass the safety guardrails built into popular AI systems, including Grok and Gemini. Called 'Cryptographic Context Injection,' the method conceals malicious instructions in an encrypted form so that AI safety filters cannot detect them. The hidden instructions are only decrypted once they reach a trusted execution environment inside the system, at which point the AI may act on them without recognising them as harmful.
This discovery highlights a growing concern for businesses adopting AI tools: safety filters designed to block harmful requests can potentially be circumvented by clever technical tricks. As more Australian small businesses use AI chatbots and assistants for customer service, content creation, and internal operations, vulnerabilities like this could be exploited to generate harmful, misleading, or unauthorised outputs.
While this is primarily a concern for the developers of these AI platforms to address through patching and improved safeguards, business owners should stay informed about the AI tools they rely on and any related security advisories. Understanding that AI safety measures are not foolproof is an important step in managing the risks of AI adoption responsibly.