Recompute or Reuse? Analyzing Stale Reasoning in Vision-Language Models
Abstract
Vision-language models can answer a question about an image correctly in isolation yet change their answer when given reasoning from an earlier image. We study which parts of this reasoning affect the answer and whether their influence continues after the model answers correctly. We call the tendency of prior reasoning to favor the old answer despite changed visual evidence a textual shortcut. Across 16 VLMs, we construct contexts containing the current image and an earlier response. We compare deleting sentences selected by fixed evidence rules, deleting non-evidence sentences, and masking explicit answer cues. Deleting the selected sentences shifts answer preference toward the current answer in all 16 models relative to non-evidence deletion; changing sentence order has a smaller and more variable effect. After an initially correct response, deleting textual evidence for the current answer causes a larger increase in old answers when the earlier reasoning is retained. When the evidence is kept, initially correct responses show little old-answer repetition, although accuracy falls on follow-up calculations. Responses that initially retain the old answer show much more repetition and more follow-up results matching calculations with the old value. These findings show which prior text affects visual updating and how this influence varies with the initial answer and the follow-up question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.