The 'Safety Penalty': Are Restrictive AI Models Slowing Down Cyber Defence?
As AI tools become more common in cybersecurity defence, a new concern is emerging: the very safety restrictions built into leading AI models may be making it harder for security teams to respond quickly to incidents. Cisco Talos describes this as a "safety penalty" — a trade-off where cautious, heavily filtered AI responses can slow down the fast decision-making needed during an active cyberattack.
The issue is that while defenders are working with AI models designed to be careful and restrained, attackers are not bound by the same limitations. Cybercriminals can use unrestricted or custom-built AI tools without safety guardrails, potentially giving them a speed advantage during real-time attacks. This creates an uneven playing field where defensive AI may lag just when speed matters most.
Talos suggests organisations should consider building toward greater "operational sovereignty" — meaning more control over how their AI tools are configured and deployed — so that defensive systems can act decisively without unnecessary delays, while still maintaining appropriate safeguards. For small and medium businesses relying on third-party AI-powered security tools, this is a reminder to understand how those tools are built and where their limitations may lie.