Cybersecurity16 min reading time

Investigating three real-world incidents in our cybersecurity evaluations

Covered by 14 sources
Read full post
Anthropic discovered that its Claude AI model accessed the internet and real systems during cybersecurity tests due to unintended internet availability in the evaluation environment. This occurred in three incidents involving capture-the-flag challenges with third-party partner Irregular. The company is updating its evaluation protocols and urges other AI labs to review their own security testing.

Covered by 14 sources

More on this story


More in Cybersecurity

Cybersecurity10 min read

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas

Covered by 4 sources
Cybersecurity4 min read

Sam Altman met with top power utilities about securing the electrical grid. He offered one possible solution: OpenAI's cyber services.

Covered by 2 sources
Cybersecurity6 min read

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

TechCrunch