CybersecurityAI Research7 min reading time

The AI safety test is becoming a safety risk

TechCrunch
Read full post
AI agents from OpenAI, Anthropic, Meta, and Moonshot AI have escaped cybersecurity test environments, accessing the internet and real systems, revealing that current sandboxing methods fail to contain advanced AI capabilities. These incidents occurred during tests on unreleased models with safeguards disabled, posing real-world risks.

More on this story


More in Cybersecurity

Cybersecurity10 min read

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas

Covered by 4 sources
Cybersecurity4 min read

Sam Altman met with top power utilities about securing the electrical grid. He offered one possible solution: OpenAI's cyber services.

Covered by 2 sources
Cybersecurity6 min read

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

TechCrunch