acceptodds
Under review as a conference paper at ICLR 2027

LoopDistill: Student-Aware Knowledge Distillation for Looped Language Models

Abstract

Looped language models (LoopLMs) increase effective depth by repeatedly applying shared Transformer blocks, offering a parameter-efficient approach to reasoning. Knowledge distillation typically transfers reasoning by training a smaller model on chain-of-thought trajectories from a stronger teacher. For a LoopLM, however, readouts at different recurrent depths can generate reasoning trajectories with different lengths and structures. A single teacher trajectory may therefore provide uneven supervision across loops, while training only the final readout leaves intermediate loops without direct reasoning targets. We introduce LoopDistill, a student-aware distillation framework that uses trajectories generated at selected depths to construct loop-specific teacher supervision. Each selected loop is trained on a target adapted to its own reasoning behavior. Using Ouro-2.6B-Base and two teachers, we evaluate LoopDistill on AIME24, AIME25, AIME26, and OlympiadBench. With Qwen3-30B-A3B as the teacher, LoopDistill outperforms the competing baseline on every benchmark, by up to 18.9 percentage points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.