A novel defensive strategy against artificial intelligence-powered cyberattacks has emerged, leveraging the very safety mechanisms that hackers often attempt to bypass.

Researchers have discovered that triggering an AI system's built-in content restrictions—often referred to as 'banned topics'—can effectively halt autonomous hacking agents in their tracks.

This approach turns the attacker's reliance on AI models into a vulnerability, using the models' own ethical guardrails as a tripwire.

The finding comes as the cybersecurity landscape grapples with the rapid evolution of AI-driven threats.

Recent incidents have highlighted the growing sophistication of these attacks.

Google's Threat Intelligence Group recently reported thwarting a planned mass exploitation operation by a hacker group that utilized artificial intelligence models to exploit previously unknown vulnerabilities.