acceptodds
Under review as a conference paper at ICLR 2027

Computation Modules for Generation: A Design Space for Label-Free Self-Distillation

Abstract

Self-distillation offers a promising route to language model self-improvement, but its effectiveness depends on the supervision a model can provide for itself. While ground-truth labels are not accessible, existing approaches often construct this supervision through heuristic generation procedures, leaving open how additional computation should be organized to make the model a more informative teacher. We introduce *Compound Computation for Amplification*, a framework for organizing self-improvement into generation and learning. At its core is a *computation module for generation*, which organizes model calls and programmatic operations as a directed acyclic graph to construct supervision from unlabeled input problems. By connecting the computation used for generation, the content it produces, and subsequent learning, the framework enables generation quality and downstream training outcomes to be studied separately. Within this design space, we compare different strategies for scaling computation and find that higher answer accuracy does not consistently lead to better self-distillation. Our observations suggest that changes in response coverage and reasoning patterns may also contribute to differences in supervision quality. Motivated by this finding, we develop *Computation-Augmented Self-Distillation (CASD)*, a concrete label-free on-policy self-distillation method within our framework, whose computation module for generation is designed to construct more effective supervision. Experiments on mathematical reasoning show that *CASD* achieves performance comparable to or better than baselines trained with ground-truth labels and outperforms all evaluated label-free baselines. These results demonstrate the potential of organizing generation-time computation to construct more effective supervision for self-improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.