acceptodds
Under review as a conference paper at ICLR 2027

Same Problem, Different Start: Controlling Text-Image Modality Gaps via First-Decision Distillation

Abstract

Recent work has shown that presenting the same problem in Image rather than Text form can induce substantial performance and reasoning differences in multimodal large language models (MLLMs), while distilling full Text-side reasoning traces can substantially recover Image-side performance. But does visual input truly impair the model's ability to reason, or does it simply send a capable model down the wrong generation trajectory, such that minimal steering can redirect it? We find that a remarkably localized intervention can suffice: altering only the first generation decision can substantially redirect the frozen continuation. On a held-out set of Image-presented GSM8K problems, suppressing answer entry only at the first generation position raises accuracy from 41.7% to 93.0%; conversely, forcing naturally non-answer-start Text responses toward answer entry reduces accuracy from 94.8% to 28.9%. Motivated by this finding, we introduce First-Decision Distillation (FDD), which distills only the teacher's first-step decision and uses a clean handoff to leave all subsequent generation to the frozen base model. With only 2 paired training examples and 16 presentations per task, FDD-Logit reduces the Qwen3-VL-8B Text–Image accuracy gap from 12.10 to 1.03 percentage points (pp), a 91.5% reduction. Evaluations across model families, presentation directions, prompt formulations, and structured-chart inputs further test the scope of first-decision control, including an improvement of 10.92 pp from reverse Image-to-Text FDD on Gemma-3-12B. These results identify response onset as a high-leverage, context-dependent control point for reducing Text–Image modality gaps without imitating the teacher's full response.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.