BETC: Bidirectional Evidence-Target Co-Inference for Reliable Multimodal Prompt Grounding
Abstract
Multimodal prompt grounding aims to identify and localize user-specified targets in an image given natural language descriptions and visual exemplars. Existing approaches typically treat prompts as reliable instructions and directly fuse the provided information. However, due to ambiguous descriptions or incomplete observations, real-world prompts may contain both informative and misleading evidence. When different evidence sources conflict, blindly aggregating all evidence can lead to unreliable grounding. We argue that reliable grounding requires more than combining multimodal information: it requires determining which evidence truly explains the target being grounded. Based on this observation, we introduce Bidirectional Evidence–Target Co-Inference (BETC), a framework that models multimodal grounding as target-relative evidence reasoning. BETC jointly reasons over evidence and targets through a bidirectional process: evidence helps identify potential targets, while target-conditioned reasoning further determines which evidence truly supports each target. We establish two complementary protocols based on PACO and PhraseCut to evaluate grounding under partially reliable prompts, covering controlled attribute perturbations and realistic prompt mismatches. Experiments demonstrate that BETC improves robustness by dynamically adjusting evidence responsibility according to target context while maintaining competitive performance on clean inputs. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.