Probabilistic Reasoning Over Visual Evidence for Compositional VQA
Abstract
Training-free compositional methods answer visual questions by chaining pretrained vision and language models to collect intermediate evidence and compose answers. Methods that execute a program to answer a question make hard decisions on each perception, discarding its uncertainty, while methods that keep perceptions uncertain follow a fixed evidence collection strategy. We introduce PROVE (Probabilistic Reasoning Over Visual Evidence), a training-free neuro-symbolic framework for compositional visual question answering that adaptively collects uncertain visual evidence. PROVE preserves perception uncertainty with calibrated probabilistic logic facts that define a Bayesian network, then answers the query by exact inference with ProbLog. Under matched components on four benchmarks, PROVE outperforms five training-free compositional methods by up to points and improves over every VLM it is built on by up to points. A deterministic ablation that rounds each perception to true or false shows that preserving uncertainty adds up to points, with the gain rising to points on questions with the most borderline perceptions. PROVE also inherits gains in LLM reasoning, as thinking mode alone adds up to points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.