acceptodds
Under review as a conference paper at ICLR 2027

Learning to See from Reading: Cross-Modal Self-Distillation for Coding

Abstract

A multimodal model can solve the same problem differently depending on whether the input is presented as text or as an image, even when both views preserve the same content. This asymmetry suggests a natural question: can the model use its stronger view to improve its weaker one? We study this question through same-model cross-modal on-policy distillation (OPD) for coding from rendered images. Specifically, an image-conditioned student generates its own responses, while a frozen copy of the same pretrained model reads the paired source text and provides next-token distributions along the student’s prefixes. Matching these distributions transfers supervision from the stronger text view without ground-truth solutions, task rewards, or a stronger external teacher. We further characterize this mechanism through projected view repair, which separates raw cross-modal disagreement from the component expressible by the student’s adaptation tangent and its alignment with task improvement. Empirically, mean image-input accuracy on MBPP+ rises from 46.6% to 52.6%, outperforming ground-truth supervised fine-tuning. Matched controls show that text-conditioned feedback yields a stronger image-student endpoint than image-conditioned feedback, while exact transcription also improves across rendering densities. The gains extend across views: image-trained checkpoints improve text-input coding beyond the frozen pretrained teacher. Together, these results show that a model can use a stronger representation of the same task as a source of self-supervision for a weaker one.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.