AI ResearchMachine Learning5 min reading time

Scale AI Reports ROK-FORTRESS Findings on Multilingual AI Safety

Unite.AI
Read full post
Scale AI and the Korea AI Safety Institute released ROK-FORTRESS, a bilingual English-Korean adversarial safety benchmark evaluating 14 frontier AI models. The study found Korean-language prompts grounded in Korean contexts led to lower measured harm across models. The benchmark assesses responses in national security and public safety domains using calibrated LLM judges and expert rubrics.

More in AI Research

AI Research3 min read

OpenAI reports 6 new instances of 'concerning model behavior' since March

Covered by 12 sources

Our framework for reporting model misalignment

Covered by 8 sources
AI Research5 min read

AI companies must work with the research community to protect attribution

Covered by 2 sources