Visual Causal Transport Distillation via Counterfactual Relative Alignment
Abstract
Visual on-policy self-distillation transfers privileged visual information mainly through a teacher's output predictions, leaving how visual evidence is used implicit. In a training-set diagnostic, image-removed accuracy rises from 35.8% to 41.6% after Vision-OPD, suggesting that part of the acquired knowledge is accessible through image-independent priors. We introduce Visual Causal Transport Distillation (\method), which augments output distillation with counterfactual relative alignment of hidden representations. For each student-generated prefix, an EMA teacher is evaluated with privileged visual input and with its visual embeddings zeroed. Their contrast provides a reference for aligning the student's final representation and cross-depth updates toward evidence-conditioned computation. A dimensionless distance ratio performs relative alignment, while intervention energy and bounded evidence-gap weights focus supervision on visually sensitive tokens and trajectories. Across six benchmarks, \method achieves average accuracies of 72.10% and 75.11% with Qwen3.5-4B and Qwen3.5-9B, outperforming Vision-OPD by 1.39 and 3.57 percentage points, respectively. \method adds only one intervened teacher pass during training and leaves student inference unchanged.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.