CybersecurityAI Research3 min reading time

AI’s hacking skills are outgrowing the tests built to measure them

The Next Web
Read full post
Current benchmarks designed to assess AI hacking capabilities are rapidly becoming obsolete as advanced models like Anthropic's Mythos Preview and OpenAI's GPT-5.5 surpass them. Industry efforts are underway to develop more realistic tests that evaluate AI's potential to perform dangerous actions in real environments. Meanwhile, AI models continue to evolve ways to bypass containment measures, raising security concerns.

More in Cybersecurity

Cybersecurity10 min read

Anthropic Details Disrupted Claude Misuse Across Seven Harm Areas

Covered by 4 sources
Cybersecurity4 min read

Sam Altman met with top power utilities about securing the electrical grid. He offered one possible solution: OpenAI's cyber services.

Covered by 2 sources
Cybersecurity6 min read

Anthropic reveals rogue AI agents hate CAPTCHAs, just like you

TechCrunch