Scaling Laws for Mixture Pretraining Under Data Constraints

Apple Research Blog
Read full post
Researchers analyzed over 2,000 language model training runs to understand how mixing scarce target data with abundant generic data affects performance. They found that repeated use of limited target data can improve results up to 15-20 repetitions, depending on model size and compute budget. They propose a scaling law to optimize data mixtures for pretraining under data constraints.

More in LLM & Text Generation

Peter Thiel-Backed AI Startup Cognition Raises Funds at $48 Billion Valuation

Covered by 2 sources

Build more natural voice experiences with GPT‑Live‑1 in the API

Covered by 2 sources