Back to feed
arXiv cs.LG·

Uncertainty-Aware LLM-Guided Policy Shaping for Sparse-Reward Reinforcement Learning

Signal
72
Hype
28
In three linesULPS integrates a calibrated LLM into the RL training loop to guide policy via A*-based symbolic trajectories and MC dropout uncertainty. On MiniGridUnlockPickup, the framework improves execution accuracy by 9%, reduces environment interactions, and increases reward AUC over unguided baselines.
Read source
Your take?
Reinforcement learningReasoningAI AgentsPapers

Summary generated by Claude — human-verified