acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Causal Distillation: Selecting Sufficient Causal Supervision

Abstract

Distilling a large language model (LLM) requires choosing which aspects of its computation a smaller student should preserve. Responses to interventions on intermediate variables and computational routes make this choice testable. We propose Hierarchical Causal Distillation (HCD), based on the principle that a distilled model should preserve a compact causal structure sufficient for the required teacher responses. HCD asks which supervision objectives enable a student to retain these responses under intervention. HCD organizes supervision around behavior, high-level reasoning programs, their circuit realizations, and internal states. It trains students with different combinations of objectives and evaluates them against common task-specific behavioral and interventional requirements. Cross-model correspondences relate semantic variables and circuit roles, allowing teachers and students with different architectures to preserve the same required responses. We formalize the choice of causal granularity using Unified Causal Minimum Description Length (UCMDL), which selects the shortest adequate joint description of a teacher mechanism and cross-model alignment within a fixed candidate family. All candidates face the same intervention requirements. A finite-family confidence bound defines adequacy on independent selection data before description-length comparison. We report results on two controlled transformers with known programs and pretrained Qwen3.5 students on seven generated arithmetic tasks. In the controlled study, HCD selects program-and-circuit supervision (BPC) on both tasks. The evaluated BPC students achieve lower ordinary and intervention response errors than behavioral KD. The independently trained frequency-sorting student also achieves low state-intervention error without a separate state objective in student training. In the pretrained comparison, BPC improves base-answer accuracy by 30.0% relative to behavioral KD and achieves the lowest pooled mean teacher–student total variation on ordinary, program-intervention, and circuit-intervention queries. These results support program and circuit supervision for preserving the evaluated teacher responses during compression.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.