SD: Skill-Conditioned Self-Distillation for LLM Social Reasoning
Abstract
While large language models (LLMs) have achieved impressive performance across many domains, their social intelligence, including the ability to reason about others’ beliefs, intentions, and emotions, remains comparatively underexplored. Recent work has applied reinforcement learning (RL) to improve LLMs’ social reasoning capabilities, but sparse verifier outcomes provide little guidance for the intermediate reasoning process. On-policy self-distillation (OPSD) alleviates this by using teacher-side privileged information to provide dense token-level supervision. However, conventional sample-specific privileged signals may fail to capture diverse, reusable reasoning strategies required for social reasoning. To address these limitations, we introduce Skill-Conditioned Self-Distillation for Social Reasoning (SD), a framework that uses structured social reasoning skills as teacher-only privileged information. We extract compact natural-language planning skills from the model’s own correct reasoning trajectories, where each skill describes a successful workflow for solving a class of social reasoning problems. During training, relevant skills are dynamically retrieved and provided only to the teacher, while the student always responds to the plain task prompt and learns to internalize the planning guidance through token-level distillation. We further combine this dense distillation signal with outcome rewards to jointly supervise both the reasoning process and the final answer. Experimental results on multiple social reasoning benchmarks demonstrate that SD substantially outperforms the standard RL baseline, improving both vanilla GRPO and vanilla OPSD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.