Faith-GVT: Evaluating Faithfulness in Grounded Visual-Textual Chain-of-Thought
Abstract
Grounded Visual-Textual Chain-of-Thought (GVT-CoT) makes multimodal reasoning more inspectable by exposing both textual rationales and explicit visual evidence regions. Yet existing faithfulness studies in vision–language models (VLMs) largely examine whether textual reasoning reflects influential cues, leaving underexplored whether textual and regional evidence claims are jointly faithful to model behavior and mutually consistent. We address this gap with Faith-GVT, a unified framework that evaluates GVT-CoT along three complementary dimensions: textual faithfulness, visual faithfulness, and text–visual consistency. Using controlled interventions across image, text, and region channels alongside predicted-region masking, we test whether reported textual and regional evidence responds to behaviorally influential information. Across 51,463 examples from five datasets and four VLMs, we find systematic dissociations among textual disclosure, reported localization, and behavioral sensitivity to interventions. Textual faithfulness averages 83.16% under text interventions but only 34.66% under misleading-region interventions, revealing a substantial gap between text and region influence disclosure. The localized bounding box is not uniformly necessary for preserving the model's answer and can diverge from the textual rationale, particularly in fine-grained document reasoning. Matched comparisons with text-only CoT further show that the effect of explicit grounding on textual faithfulness is model- and intervention-dependent. These results show that grounding introduces additional evidence claims whose faithfulness cannot be assessed by text-only methods alone: textual disclosure, reported localization, behavioral sensitivity, and text–visual consistency capture distinct aspects of multimodal reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.