acceptodds
Under review as a conference paper at ICLR 2027

Where to Go, What to Avoid: State-Relative Self-Distillation

Abstract

Existing on-policy self-distillation learns from the student's own rollouts, but its teacher signal remains conditioned on a fixed reference solution. This can entangle useful local reasoning with answer-dependent confidence and preference for a particular solution path. We argue that supervision should instead reflect the student's current rollout state, encouraging viable continuations while discouraging directions that have already led to failure. We therefore propose **State-Relative Self-Distillation (SRSD)**, which incorporates state-relative rollout information into supervision through two complementary teacher roles. A State Tracker constructs a state representation from the student's complete on-policy rollout, providing a positive *Where to go* signal for subsequent reasoning. A State Guard leverages a failed trajectory sampled from the same policy to provide a negative *What to avoid* signal, discouraging the student from revisiting failure-inducing directions. Together, these signals encourage valid alternative reasoning paths while reducing unsupported confidence and rigid path preference. We evaluate SRSD on mathematical reasoning, code generation, and agentic tool use across multiple model scales and benchmarks. SRSD consistently delivers strong performance improvements across these domains, with the largest gains in code generation, where it outperforms the strongest compared method by 8 points. Code will be released upon publication to facilitate reproducibility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.