Back to feed
arXiv cs.LG·

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

Signal
72
Hype
18
In three linesSA-AH-GRPO, an extension of GRPO, applies asymmetric entropy-based token discounting for LLM reinforcement learning. On GSM8K, Qwen 2.5-3B achieves Pass@1=0.858 with 3.6× variance reduction vs standard GRPO while preserving gradients on correct trajectories.
Read source
Your take?
Reinforcement learningReasoningBenchmarks

Summary generated by Claude — human-verified