DevMachine Learning10 min reading time

The LLM Judge That Kept Agreeing With Itself

Towards Data Science
Read full post
A system using a language model as a judge to approve SQL queries repeatedly approved incorrect queries due to structural bias, prompting developers to treat the judge as a component needing its own testing and calibration against human reviewers.

More in Dev

Introducing the Agents API

Covered by 3 sources
Dev1 min read

Datasette 1.0a39 and 0.65.4 security releases

Simon Willison's Weblog
Dev1 min read

Native is now the future of mobile at Shopify

Simon Willison's Weblog