Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models
Signal
72
Hype
18
In three linesSA-AH-GRPO, an extension of GRPO, applies asymmetric entropy-based token discounting for LLM reinforcement learning. On GSM8K, Qwen 2.5-3B achieves Pass@1=0.858 with 3.6× variance reduction vs standard GRPO while preserving gradients on correct trajectories.Read source
Your take?
Summary generated by Claude — human-verified