Machine LearningAI Research2 min reading time

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Apple Research Blog
Read full post
Researchers developed a rubric-based reward system for open-domain question answering that uses query-specific, evidence-grounded rubrics decomposed into multiple quality dimensions. This approach improves answer quality across composition, grounding, and instruction-following by up to 6.5%. Conditioning rubrics on retrieved evidence enhances factual accuracy, while multi-dimensional rubrics improve coherence and adherence to query requirements.

More in Machine Learning

Machine Learning3 min read

Anthropic caught scientists using Claude to further biological weapon research

Covered by 2 sources
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

Covered by 2 sources
Machine Learning4 min read

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

The New Stack (AI)