Behavioral Traits as Privileged Supervision for Self-Distilled Reasoning
Abstract
Behavioral traits are commonly studied as alignment targets for large language models. These general dispositions also shape how models approach reasoning and problem solving, motivating their use as reusable teaching principles. We propose trait-guided on-policy self-distillation (-OPSD), which converts behavioral traits into privileged supervision for reasoning. A trait-conditioned self-teacher provides dense token-level guidance on student-generated trajectories, enabling the student to learn from these qualities without retaining trait instructions at inference. We instantiate the framework with truthfulness, encouraging justified inference, uncertainty checking, and self-correction. To adapt this guidance, -OPSD uses hard-sample learning to provide local-progress instructions for budget-exhausted trajectories and an outcome-aware stance gate to reinforce, suppress, or exclude supervision according to its consistency with task outcomes. Reference answers are used only for verification and never enter the teacher context. Across Qwen3 models from 1.7B to 8B, -OPSD achieves the highest average mathematical accuracy among the compared methods and improves abstention under insufficient evidence. Gains with additional behavioral traits, including planning and exploration, demonstrate that the framework generalizes beyond truthfulness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.