AgentsMachine Learning9 min reading time

Claude performed best on a new benchmark for ‘agents that build agents’. But it passed fewer than a quarter of the tests.

The New Stack (AI)
Read full post
Sierra's Hyper-𝜏-bench benchmark evaluates how well AI developer agents can autonomously build other agents. Testing six model and coding harness combinations, Claude Opus 5 achieved the highest success but still passed fewer than 25% of tasks.

More in Agents

Meta Announces Muse AI Agent for Personal Tasks and Organization

Covered by 11 sources

Introducing the Agents API

Covered by 3 sources
Agents4 min read

Amazon makes its agentic AI platform Quick generally available for desktop on Windows and macOS

Covered by 2 sources