Back to feed
arXiv cs.CL·

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

Signal
72
Hype
35
In three linesOPD-Evolver is a slow-fast co-evolution framework that cultivates self-evolving agents through on-policy self-distillation. The system manages a four-level memory hierarchy to read, use, write, and maintain experience. Across multi-domain benchmarks, OPD-Evolver outperforms ReasoningBank (+11.5%) and Skill0 (+5.8%), with OPD-Evolver-9B rivaling Qwen3.5-397B and Step-3.5-Flash.
Read source
Your take?
AI AgentsReasoningReinforcement learningBenchmarks

Summary generated by Claude — human-verified