OpenAI’s AGI number came from a harness, not the model
The Next Web
Read full postOpenAI's GPT-6 Astra achieved a 99.9% AGI benchmark score using its proprietary Provider Adapter harness, but the same model scored 62.7% under the standard ARC Prize harness. The difference stems from the software environment around the model, not the model itself, highlighting the impact of evaluation setup on AGI claims.

- OpenAI spent millions to solve this famous math problem — mathematicians are furious· Understanding AI
- OpenAI puts Pro subscriptions on hold due to Astra demand· TechCrunch
- OpenAI’s Supposed Mathematical Breakthrough Devolves Into Explosive Drama as Mathematician Accuses It of Stealing His Work· Futurism
- OpenAI Releases GPT-6 Astra for Coding and Computer Use· InfoQ (AI, ML & Data)
- EU cybersecurity agency is now testing Mythos 5 and GPT-6 Astra, the Commission says· The Next Web
- The Fight Over OpenAI’s Math Breakthrough Is a New Kind of Scientific Arms Race· 2 sources


