Truth in Focus: Scene Graph Grounding for Faithful Visual Question Answering
Abstract
Recent advances in large vision–language models (VLMs) have substantially improved Visual Question Answering. Nevertheless, reliable answer generation remains a persistent challenge, as VLMs may produce plausible responses that are insufficiently grounded in the visual content. Scene graphs provide a structured representation of visual content by organizing object entities and their relationships. However, a complete scene graph often contains information beyond what a given question requires, which can distract the model from relevant evidence and lead to factually unsupported answers. We therefore shift from global scene modeling to question-conditioned local grounding and propose a scene graph grounding framework that extracts a compact subgraph containing the visual information needed to answer each question. Specifically, we introduce a dual-differential causal attention module that jointly models node–node dependencies and node–query associations, enabling the relative contribution of graph elements to be assessed for subgraph prediction. To mitigate the scarcity of fine-grained grounding supervision, we further construct a factual–counterfactual dataset by leveraging the generative capabilities of large models. Extensive experiments on multiple benchmarks demonstrate that our method achieves superior performance against state-of-the-art approaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.