DOES SELF-TRAINING ON PRIVILEGED REPAIRS CLOSE REPRESENTATION GAPS? A DIAGNOSTIC EVALUATION IN FOUR DOMAINS
Abstract
Vision-language models often answer a question correctly from one representation of an object and incorrectly from an informationally equivalent one. We evaluate a strategy that turns such disagreements into supervision: the model, shown both representations and both of its traces, rewrites its failed trace, and a LoRA student that sees only the weak representation is trained to match the frozen model’s token distribution over the rewrite, alongside replay of questions it already answers correctly. The privileged context thus decides what the student is trained toward. Using Qwen3-VL-8B-Instruct across four domains (synthetic social graphs, Lichess positions, WikiTableQuestions and MolLangBench), we measure what the strategy changes relative to the frozen model and where those changes come from. It narrows the representation gap sharply on chess and graphs, where the weak representation is the image (the chess image arm gains 17 points), gives small gains on tables, and widens the gap on molecules, where the weakest representation barely moves; the representations the model already reads well change little. The gains lean toward questions that point to a specific part of the image, such as one node’s arrows or one row of the board, over questions that require searching or comparing across the whole picture, with exceptions in both directions. Against answer-only fine-tuning on the same data the results are mixed; it sometimes transfers better. Rerun from its own trained model, it improves once more and then stops. The strategy cannot supervise a representation its teacher cannot reason in, needs two additions to train stably, and can leave correct answers resting on reasoning that contradicts the data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.