acceptodds
Under review as a conference paper at ICLR 2027

SocraticPO: Policy Optimization via Interactive Guidance

Abstract

Outcome-based reinforcement learning (RL) improves large language model reasoning by reinforcing correct responses, but leaves failed attempts without a path to correction. Inspired by Socratic teaching, in which a human teacher uses questions and hints to prompt learners to examine and repair their own mistakes, we introduce SocraticPO (Socratic Policy Optimization) to model failed reasoning as a teacher–student correction process. After an independent attempt fails, the teacher provides guidance conditioned on that attempt, and the student generates its own revision, shifting learning from outcome selection to guided correction. Rewarding assisted and independent solutions equally can encourage reliance on the teacher. SocraticPO therefore uses assistance-aware, batch-adaptive reward decay to favor solutions requiring less guidance while preserving non-negative normalized advantages for successful corrections. Across four SciKnowEval scientific reasoning domains and GSM8K, SocraticPO outperforms conventional RL and self-distillation baselines in unassisted accuracy under a controllable additional inference budget. Further analyses show that a teacher's answer accuracy does not necessarily predict its guidance quality and that the gains do not require a consistent reduction in teacher–student KL divergence, highlighting the advantage of text-level guidance over answer or distribution transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.