acceptodds
Under review as a conference paper at ICLR 2027

Behavioral Traits as Privileged Supervision for Self-Distilled Reasoning

Abstract

Behavioral traits are commonly studied as alignment targets for large language models. These general dispositions also shape how models approach reasoning and problem solving, motivating their use as reusable teaching principles. We propose trait-guided on-policy self-distillation (-OPSD), which converts behavioral traits into privileged supervision for reasoning. A trait-conditioned self-teacher provides dense token-level guidance on student-generated trajectories, enabling the student to learn from these qualities without retaining trait instructions at inference. We instantiate the framework with truthfulness, encouraging justified inference, uncertainty checking, and self-correction. To adapt this guidance, -OPSD uses hard-sample learning to provide local-progress instructions for budget-exhausted trajectories and an outcome-aware stance gate to reinforce, suppress, or exclude supervision according to its consistency with task outcomes. Reference answers are used only for verification and never enter the teacher context. Across Qwen3 models from 1.7B to 8B, -OPSD achieves the highest average mathematical accuracy among the compared methods and improves abstention under insufficient evidence. Gains with additional behavioral traits, including planning and exploration, demonstrate that the framework generalizes beyond truthfulness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.