Different Semantic Edits Converge to Grounded Truth in Vision-Language Models
Abstract
Vision-language models (VLMs) must determine how a semantic change alters the truth of a proposition given visual evidence, rather than respond to surface cues such as negation or object identity. Yet how this grounded truth computation emerges inside a VLM remains unclear. We study this question using answer-balanced natural-image quartets in which negation and object substitution each reverse the correct answer in both directions. Across three VLMs and two natural-image benchmarks, we identify a common edit-to-truth transition: early representations preserve edit-specific structure, whereas intermediate layers organize distinct edits according to the truth change required by the image. This shared representation is both causally answer-relevant and visually grounded. At the transition, a one-dimensional truth-axis projection retains over 80% of the full attention-swap effect for both edits, while the orthogonal residual has little effect; the same geometric progression also appears when only the visual evidence changes. Localized removal of visual evidence suppresses the truth representation and its causal effect but largely preserves linguistic polarity. The resulting representation follows proposition truth rather than the literal output symbol under response-code reversal, and frozen RePOPE→AMBER probes transfer without target fitting, reaching 91–93% accuracy over all target images. Finally, source-derived truth directions selectively repair naturally occurring errors: on LLaVA, they correct 103 AMBER errors without introducing new ones, raising accuracy from 79.0% to 87.6%. Together, these results reveal a shared grounded-truth stage through which distinct semantic edits are transformed into visually grounded decisions in VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.