SocraticPO: Policy Optimization via Interactive Guidance
SocraticPO is a policy optimization framework for LLMs that augments RL rollouts with Socratic-style natural-language guidance. A teacher diagnoses errors and provides concise corrections, paired with reward decay to prevent the model from relying on assistance. Tested on SciKnowEval, it outperforms RL and self-distillation baselines.