PAVER: Dual-Source Parsing and Premise-Aware Routing for Reliable LVLM Responses
Abstract
Large vision-language models (LVLMs) can produce unreliable responses for two distinct reasons: a user question may presuppose visual conditions that do not hold, while the model may introduce unsupported content in its answer. These failures require different treatment because a visual proposition may be a required question premise, queried content, or a model assertion. We introduce PAVER (Premise- and Answer-aware Visual Evidence Routing), a training-free framework that operationalizes this distinction asymmetrically. Required question premises determine whether the question remains valid, whereas entities exposed by an initial draft only guide targeted visual inspection. Grounding DINO and Florence-2 provide shared visual evidence for premise-aware routing and final response generation. The router selects refusal when a required premise is contradicted and otherwise selects evidence-grounded re-answering. Across six benchmarks and two LVLM backbones, PAVER reduces answer-side hallucination and improves premise-sensitive reliability. Factorized ablations further show complementary roles: the draft-guided pathway improves answer-side reliability, while explicit premise modeling improves premise-sensitive behavior. These results support interaction-role-aware evidence routing as a practical principle for reliable LVLM responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.