Same Problem, Different Start: Controlling Text-Image Modality Gaps via First-Decision Distillation
Abstract
Recent work has shown that presenting the same problem in Image rather than Text form can induce substantial performance and reasoning differences in multimodal large language models (MLLMs), while distilling full Text-side reasoning traces can substantially recover Image-side performance. But does visual input truly impair the model's ability to reason, or does it simply send a capable model down the wrong generation trajectory, such that minimal steering can redirect it? We find that a remarkably localized intervention can suffice: altering only the first generation decision can substantially redirect the frozen continuation. On a held-out set of Image-presented GSM8K problems, suppressing answer entry only at the first generation position raises accuracy from 41.7% to 93.0%; conversely, forcing naturally non-answer-start Text responses toward answer entry reduces accuracy from 94.8% to 28.9%. Motivated by this finding, we introduce First-Decision Distillation (FDD), which distills only the teacher's first-step decision and uses a clean handoff to leave all subsequent generation to the frozen base model. With only 2 paired training examples and 16 presentations per task, FDD-Logit reduces the Qwen3-VL-8B Text–Image accuracy gap from 12.10 to 1.03 percentage points (pp), a 91.5% reduction. Evaluations across model families, presentation directions, prompt formulations, and structured-chart inputs further test the scope of first-decision control, including an improvement of 10.92 pp from reverse Image-to-Text FDD on Gemma-3-12B. These results identify response onset as a high-leverage, context-dependent control point for reducing Text–Image modality gaps without imitating the teacher's full response.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.