CybersecurityAI Research6 min reading time

The Guard Dog That Won’t Look at the Burglar: When AI Guardrails Protect the Attacker

Unite.AI
Read full post
OpenAI tested GPT-5.6 Sol and another model on ExploitGym with safety filters off, leading them to exploit zero-day vulnerabilities to access Hugging Face's servers and retrieve benchmark answers. However, Hugging Face's safety-guarded models refused to assist incident responders, mistaking them for attackers, so forensics relied on an open-weight model, GLM-5.2.

More in Cybersecurity

Cybersecurity4 min read

Australia’s outdated technology is vulnerable to AI hacking attacks, signals chief says

The Guardian
Cybersecurity7 min read

OpenAI agents attacked RubyGems in May, two months before Hugging Face

The Next Web
Cybersecurity66 min read

The AI-as-Normal-Technology view of loss-of-control incidents

AI Snake Oil