acceptodds
Under review as a conference paper at ICLR 2027

RegionTrace: Measuring How Image Regions Shape Multimodal Generation

Abstract

Reliable multimodal assistance requires understanding whether answers are grounded in relevant visual evidence. We present RegionTrace, a framework for tracking how image regions support the same answer throughout generation. At each stage, it holds the text generated so far fixed and measures how restoring an image region changes the answer score. Keeping the answer and regions fixed reveals where attribute errors find support, how support evolves during explanation, and whether visual revisits restore it. Experiments across three models and four benchmarks reveal that models often borrow attributes from the wrong object: in color-misbinding errors, the color-bearing distractor outweighs the queried region in 68.0–97.9 % of cases. Across 7,206 explanations, a complete-image probe matches the eventual answer in 86–92 % of traces before explanation begins, while regional support weakens after the first sentence in all six model–task settings. Explicitly revisiting the image does not reliably restore this support or improve accuracy. RegionTrace connects spatial attribution with conditional answer-support trajectories throughout multimodal generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.