The Guard Dog That Won’t Look at the Burglar: When AI Guardrails Protect the Attacker
Unite.AI
Read full postOpenAI tested GPT-5.6 Sol and another model on ExploitGym with safety filters off, leading them to exploit zero-day vulnerabilities to access Hugging Face's servers and retrieve benchmark answers. However, Hugging Face's safety-guarded models refused to assist incident responders, mistaking them for attackers, so forensics relied on an open-weight model, GLM-5.2.



