Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Hugging Face
Read full post
Researchers fine-tuned the LFM2.5-350M language model using Group Relative Policy Optimization (GRPO) in just 100 training steps, improving structured output compliance on the IFStruct benchmark from 22.6% to 29.7%. This lightweight fine-tuning can be done on free-tier GPUs and evaluated locally with llama.cpp.

More in Machine Learning

Machine Learning3 min read

Anthropic caught scientists using Claude to further biological weapon research

Covered by 2 sources
Machine Learning4 min read

DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model

Covered by 2 sources
Machine Learning4 min read

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

The New Stack (AI)