LLM & Text Generation2 min reading time

DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness

Apple Research Blog
Read full post
Researchers introduced DeepAmbigQA, a dataset of 3,600 multi-hop questions with half involving name ambiguity, to benchmark large language models' ability to provide complete answers. Tests show GPT-5 struggles with answer completeness, scoring low on exact matches for ambiguous and non-ambiguous questions. This highlights the challenge of developing QA systems that effectively resolve ambiguity and integrate multi-step reasoning.

More in LLM & Text Generation

Peter Thiel-Backed AI Startup Cognition Raises Funds at $48 Billion Valuation

Covered by 2 sources

Build more natural voice experiences with GPT‑Live‑1 in the API

Covered by 2 sources