Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning
Signal
72
Hype
28
In three linesULPS integrates a calibrated LLM into the RL training loop to guide policy via A*-based symbolic trajectories and MC dropout uncertainty. On MiniGridUnlockPickup, the framework improves execution accuracy by 9%, reduces environment interactions, and increases reward AUC over unguided baselines.Read source
Your take?
Summary generated by Claude — human-verified