acceptodds
Under review as a conference paper at ICLR 2027

BETC: Bidirectional Evidence-Target Co-Inference for Reliable Multimodal Prompt Grounding

Abstract

Multimodal prompt grounding aims to identify and localize user-specified targets in an image given natural language descriptions and visual exemplars. Existing approaches typically treat prompts as reliable instructions and directly fuse the provided information. However, due to ambiguous descriptions or incomplete observations, real-world prompts may contain both informative and misleading evidence. When different evidence sources conflict, blindly aggregating all evidence can lead to unreliable grounding. We argue that reliable grounding requires more than combining multimodal information: it requires determining which evidence truly explains the target being grounded. Based on this observation, we introduce Bidirectional Evidence–Target Co-Inference (BETC), a framework that models multimodal grounding as target-relative evidence reasoning. BETC jointly reasons over evidence and targets through a bidirectional process: evidence helps identify potential targets, while target-conditioned reasoning further determines which evidence truly supports each target. We establish two complementary protocols based on PACO and PhraseCut to evaluate grounding under partially reliable prompts, covering controlled attribute perturbations and realistic prompt mismatches. Experiments demonstrate that BETC improves robustness by dynamically adjusting evidence responsibility according to target context while maintaining competitive performance on clean inputs. Code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.