Cybersecurity Research

Study Finds AI Chatbot 'Safety Filters' Are Surprisingly Thin — Here's Why It Matters for Your Business

Unit 42 · 29 Aug 2026
Key Takeaway Don't rely solely on an AI vendor's built-in safety features — add your own monitoring, usage policies, and output checks when deploying AI tools in your business.

Businesses increasingly rely on AI chatbots and large language models (LLMs) for customer service, content generation, and internal tools. New research from Unit 42 examined how these systems refuse unsafe or harmful requests, and found that this 'refusal' behaviour is not deeply embedded throughout the AI's reasoning — it lives in a surprisingly thin layer of the model. Small, targeted changes to how the model processes information can disrupt this layer, potentially causing the AI to bypass its own safety guardrails.

This matters because many businesses treat built-in AI safety features as a complete solution, assuming the vendor's safeguards are enough to prevent misuse, offensive outputs, or manipulation. The research suggests this confidence may be misplaced — if safety behaviour is fragile and concentrated in one narrow part of the model, it can potentially be undermined more easily than expected, whether through deliberate attacks or unexpected edge cases.

For Australian SMBs using AI-powered tools, chatbots, or third-party AI integrations, this is a reminder that internal AI safety features shouldn't be the only line of defence. Businesses should treat AI safety as one layer within a broader security strategy, including monitoring outputs, setting usage policies, and applying additional filtering or oversight controls rather than relying solely on the AI provider's built-in protections.

Building or buying AI systems? Governing them under ISO 42001 ->

Summarised by CISO AI from Unit 42. We link back to every original so you can read it yourself.