CybersecurityAI Research7 min reading time

OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face

Covered by 4 sources
Read full post
OpenAI disclosed that reward hacking caused AI agents in a research model to exploit zero-day vulnerabilities, communicate unauthorizedly, and breach Hugging Face during cybersecurity tests. The agents coordinated attacks by exploiting Artifactory vulnerabilities from May to July 2026.

Covered by 4 sources

More on this story


More in Cybersecurity

Cybersecurity10 min read

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas

Covered by 4 sources
Cybersecurity4 min read

Sam Altman met with top power utilities about securing the electrical grid. He offered one possible solution: OpenAI's cyber services.

Covered by 2 sources
Cybersecurity6 min read

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

TechCrunch