Source References Fail to Constrain Inference: Diagnosing Source Misbinding in Multi-Image Contexts
Abstract
In multi-image question answering, even when the question explicitly specifies the target image, the model may still answer based on other images in the context. We define this failure as source misbinding: the model answers correctly from the target image when it is presented alone but produces an answer consistent with a distractor source when other images are introduced. This discrepancy leads us to hypothesize that the model may fail to translate the specified source into a constraint on evidence use during inference. To test this hypothesis, we construct D-SrcBench and T-SrcBench and conduct source readout and intervention experiments. The results support the explanation that source information remains decodable within the model, yet it does not reliably constrain the evidence used for answering. Based on this diagnosis, we propose Source-First Inference: with the backbone frozen, the model first predicts the relevant image source and then downweights or masks attention to distractor-image tokens at sensitive layers. On the fixed confirmatory test set, internal hard masking (Hard) improves LLaVA-1.5-7B accuracy from 30.4% to 59.0%, outperforming soft biasing (Soft) with the same source reader (51.2%) and the best-performing evaluated baseline, our reproduction of MIA-DPO (36.0%).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.