Dev14 min reading time

Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker

Import AI
Read full post
Epoch and METR released MirrorCode, a benchmark to test AI on long-horizon programming tasks without source code access. AI models like Opus 4.7 and GPT-5.5 successfully reimplemented large programs, showing rapid improvement but some tasks remain unsolved.

More in Dev

Introducing the Agents API

Covered by 3 sources
Dev1 min read

Native is now the future of mobile at Shopify

Simon Willison's Weblog
Dev5 min read

AWS open-sources Pizza Bot: email-style inbox for background AI agents

The New Stack (AI)