Where Cross-Modal Reasoning Distillation Hurts, and How to Undo It
Abstract
Distilling chain-of-thought traces from a text-only language model into a vision-language model is a standard route to multimodal reasoning, and it carries a known side effect: the student sees worse and hallucinates more. The field prices this as the cost of the reasoning gained. But the two instruments behind that price are bent: the answer extractor misses a concise model's answers, in favor of the model that writes more, and the hallucination matcher mistakes similarity for identity and under-counts hallucination. Repaired, they show no measurable gain for the student over the base it started from on any of five benchmarks, while the cost is paid in full. That cost is not one thing but three, and each has its own remedy. (i) Answers the student never states: forcing a conclusion at inference brings them back, and no weight changes. (ii) Hallucination: it is mostly the cost of length, not of reasoning, and a teacher with shorter traces removes it. (iii) Perception cost: it has an address, resolvable only on questions that cannot be answered without the picture — a blind teacher takes the sight its student already had. Averaging the distilled weights back toward the original ones, with no training, recovers most of the lost accuracy while keeping most of the reasoning habit. We call it behavior-preserving interpolation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.