Your Agent Aced the Task. Will It Do It Again?
Hugging Face
Read full postA GPT-4.1-based ReAct agent on AppWorld shows a 24.4-point gap between average success rate (77.4%) and consistent success across repeated runs (53.0%). A new Consistency Analyzer diagnostic identifies unstable decision points, and applying consistency guidelines reduces this gap by half without lowering average accuracy.




