Back to feed
arXiv cs.LG·

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

Signal
75
Hype
25
In three linesTD-Grokking recursively decomposes zero-reward problems into verifiable subproblems to generate training signals. Tested on mathematical and medical tasks, the framework outperforms vanilla GRPO and converts zero-reward examples into usable optimization signals.
Read source
Your take?
Reinforcement learningReasoningPapers

Summary generated by Claude — human-verified