acceptodds
Under review as a conference paper at ICLR 2027

EAVI: Evidence-Aware Cross-Modal Alignment with Value-Guided Interaction for Chat-based Person Retrieval

Abstract

Chat-based person retrieval (ChatPR) identifies a target pedestrian in a gallery by progressively acquiring appearance evidence through multi-round question answering. This setting creates two requirements: a retriever must make effective use of evidence accumulated across turns, and a controller must acquire additional evidence selectively under a limited budget. We propose EAVI to address both. Its Dialogue Evidence Set Retriever (DESR) complements global matching with anonymous evidence sets aligned by entropy-regularized optimal transport and reliability-weighted aggregation, while keeping gallery representations cacheable. Its Value-Guided Interaction Agent (VGIA) selects from the recorded question pool using only the observable retrieval state, predicting each candidate's one-step retrieval value and deciding when to stop with an independent head. Because ranking questions within a dialogue and comparing values across dialogues rely on different properties of the value, VGIA treats question ordering, stopping, and shared-budget allocation as separate decisions. On ChatPedes, DESR improves R@1 by 1.7–2.4 points over the released DiaNA checkpoint under identical dialogue prefixes, with 3.0× fewer loaded parameters and about 6× faster full-test evaluation. A fitted question prior remains competitive for within-dialogue ordering, whereas retrieval-state feedback improves shared-budget allocation by 5.01 R@1 points on average over a question-and progress control across four budgets on a fixed acquisition path. A development-selected stopping head reduces additional QA acquisitions by 37.7% at a 0.67-point R@1 cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.