Back to feed
arXiv cs.CL·

Retrospective Progress-Aware Self-Refinement for LLM Agent Training

Signal
78
Hype
25
In three linesRePro, a training framework for LLM agents, teaches models to retrospectively self-generate progress signals via a forward-then-reflect rollout paradigm. Tested on WebShop, ALFWorld, and Sokoban with Qwen family, RePro achieves up to 12% absolute success rate gains without continuous external supervision.
Read source
Your take?
AI AgentsReinforcement learningReasoningQwen

Summary generated by Claude — human-verified