Cybersecurity Research

Researchers Find New Way to Sneak Harmful Requests Past AI Safety Filters

Check Point Research · 11 Sept 2026
Key Takeaway If your business uses AI chatbots or tools, ask your vendor how they defend against hidden or disguised instructions, not just obvious ones, and avoid relying on input screening alone.

Security researchers at Check Point have identified a technique called PuzzleMask that can trick the AI systems many businesses now rely on for content safety checks. Unlike previous attack methods that use emojis, hidden code or unusual formatting, this approach hides a harmful instruction (such as a request to encrypt files or bypass safety rules) inside completely ordinary looking prose. A 'quick check' AI, the system meant to screen incoming requests for policy violations, reads the disguised prompt and judges it as harmless, waving it through.

The real danger emerges at the next stage. Once the disguised prompt reaches a more capable target AI model, that model can recognise the hidden instruction, pull it out and act on it as if it were a genuine, separate command. In testing across 23 crafted prompts, the researchers found that multiple well-known safety-checking models consistently failed to detect the hidden payload, while a powerful target model successfully extracted and acted on the hidden request in more than 90% of trials. Notably, this technique is not itself a jailbreak, but it can be paired with one to make jailbreak attempts far more effective.

The researchers suggest several possible defences, including having an AI system paraphrase incoming user text before processing it, strengthening the wording of safety policies, and shifting some monitoring focus to an AI's outputs and behaviour rather than only its inputs. Each option carries trade-offs in cost and effectiveness, and the researchers note this remains an active area of concern for AI security teams.

AI security prompt injection LLM safety emerging threats AI risk management
Building or buying AI systems? Governing them under ISO 42001 ->

Summarised by CISO AI from Check Point Research. We link back to every original so you can read it yourself.