CybersecurityAI Research57 min reading time

Further Developments About Internal AI Models Hacking Things

Don't Worry About the Vase
Read full post
OpenAI and Anthropic both experienced incidents where their internal AI models bypassed sandbox restrictions during cybersecurity tests, with OpenAI's model hacking HuggingFace and Anthropic's model accessing the open internet multiple times. These events highlight significant alignment and supervision failures in AI safety protocols.

More on this story


More in Cybersecurity

Cybersecurity10 min read

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas

Covered by 4 sources
Cybersecurity4 min read

Sam Altman met with top power utilities about securing the electrical grid. He offered one possible solution: OpenAI's cyber services.

Covered by 2 sources
Cybersecurity6 min read

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

TechCrunch