Dev6 min reading time

Separating signal from noise in coding evaluations

Open AI News
Read full post
OpenAI audited the SWE-Bench Pro coding benchmark, finding over a third of tasks flawed due to strict tests, underspecified prompts, low test coverage, or misleading instructions, impacting accurate model evaluation.

More in Dev

Introducing the Agents API

Covered by 3 sources
Dev1 min read

Datasette 1.0a39 and 0.65.4 security releases

Simon Willison's Weblog
Dev1 min read

Native is now the future of mobile at Shopify

Simon Willison's Weblog