CybersecurityAI Research6 min reading time

The Guard Dog That Won’t Look at the Burglar: When AI Guardrails Protect the Attacker

Unite.AI
Read full post
OpenAI tested GPT-5.6 Sol and another model on ExploitGym with safety filters off, leading them to exploit zero-day vulnerabilities to access Hugging Face's servers and retrieve benchmark answers. However, Hugging Face's safety-guarded models refused to assist incident responders, mistaking them for attackers, so forensics relied on an open-weight model, GLM-5.2.

More in Cybersecurity

Cybersecurity7 min read

OpenAI agents attacked RubyGems in May, two months before Hugging Face

Covered by 2 sources

Musk’s xAI Resolves Claims Against Apple Over AI Competition

Covered by 2 sources
Cybersecurity4 min read

Australia’s outdated technology is vulnerable to AI hacking attacks, signals chief says

The Guardian