Learning to See from Reading: Cross-Modal Self-Distillation for Coding
Abstract
A multimodal model can solve the same problem differently depending on whether the input is presented as text or as an image, even when both views preserve the same content. This asymmetry suggests a natural question: can the model use its stronger view to improve its weaker one? We study this question through same-model cross-modal on-policy distillation (OPD) for coding from rendered images. Specifically, an image-conditioned student generates its own responses, while a frozen copy of the same pretrained model reads the paired source text and provides next-token distributions along the student’s prefixes. Matching these distributions transfers supervision from the stronger text view without ground-truth solutions, task rewards, or a stronger external teacher. We further characterize this mechanism through projected view repair, which separates raw cross-modal disagreement from the component expressible by the student’s adaptation tangent and its alignment with task improvement. Empirically, mean image-input accuracy on MBPP+ rises from 46.6% to 52.6%, outperforming ground-truth supervised fine-tuning. Matched controls show that text-conditioned feedback yields a stronger image-student endpoint than image-conditioned feedback, while exact transcription also improves across rendering densities. The gains extend across views: image-trained checkpoints improve text-input coding beyond the frozen pretrained teacher. Together, these results show that a model can use a stronger representation of the same task as a source of self-supervision for a weaker one.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.