Back to feed
arXiv cs.AI·

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs

Signal
72
Hype
45
In three linesE³RL, a reinforcement learning method, addresses error propagation in long-horizon reasoning of LLMs. Using autoregressive cross-entropy as an epistemic uncertainty signal, the model can locally correct logical defects and reuse KV cache. On AIME, 4B and 8B models outperform SOTA by 5.349% and 6.514%.
Read source
Your take?
Reinforcement learningReasoningBenchmarks

Summary generated by Claude — human-verified