acceptodds
Under review as a conference paper at ICLR 2027

RIVAD: Reasoning with Information-Variant Views for Adaptive Multi-Mode Distillation

Abstract

Large language models (LLMs) exhibit strong multi-step reasoning capabilities, yet transferring these abilities to compact models remains challenging. Existing methods typically rely on either teacher-generated trajectories, which provide stable supervision but may not reflect the prefixes encountered by the student at inference, or student-generated trajectories, which reduce this mismatch at higher training cost. We introduce RIVAD, a reasoning distillation framework that combines information-variant self-distillation with adaptive multi-mode training. Training integrates three complementary modes: canonical off-policy distillation from an external teacher, privileged-context self-distillation, and on-policy distillation over fresh student-generated responses. In the self-distillation mode, RIVAD constructs information-variant views by exposing a self-teacher to different partial, order-preserving subsets of the canonical reasoning, while the student receives only the original question. An adaptive training strategy dynamically adjusts the contribution of different supervision modes according to the student's learning progress, allowing the training allocation to evolve throughout optimization. Together, these components combine stable canonical targets, information-variant privileged guidance, and feedback on student-visited prefixes in a unified framework for distilling multi-step reasoning into compact language models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.