CladBench – an open benchmark for AI on UK building regulations

Hacker News
Read full post
CladBench is an open benchmark evaluating large language models on UK and EU building regulations with 536 questions across 12 categories. Seven models were tested, with Claude Opus 4.7 scoring highest and Gemini 2.5 Pro and GPT-5 performing comparably. The benchmark highlights varying model performance depending on task category rather than overall score.

More in LLM & Text Generation

Peter Thiel-Backed AI Startup Cognition Raises Funds at $48 Billion Valuation

Covered by 2 sources

Build more natural voice experiences with GPT‑Live‑1 in the API

Covered by 2 sources