RegionTrace: Measuring How Image Regions Shape Multimodal Generation
Abstract
Reliable multimodal assistance requires understanding whether answers are grounded in relevant visual evidence. We present RegionTrace, a framework for tracking how image regions support the same answer throughout generation. At each stage, it holds the text generated so far fixed and measures how restoring an image region changes the answer score. Keeping the answer and regions fixed reveals where attribute errors find support, how support evolves during explanation, and whether visual revisits restore it. Experiments across three models and four benchmarks reveal that models often borrow attributes from the wrong object: in color-misbinding errors, the color-bearing distractor outweighs the queried region in 68.0–97.9 % of cases. Across 7,206 explanations, a complete-image probe matches the eventual answer in 86–92 % of traces before explanation begins, while regional support weakens after the first sentence in all six model–task settings. Explicitly revisiting the image does not reliably restore this support or improve accuracy. RegionTrace connects spatial attribution with conditional answer-support trajectories throughout multimodal generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.