Distributed Local Evidence Allocation under Strong Global Anchors for Composed Image Retrieval
Abstract
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a modification text. Recent supervised CIR methods complement strong vision-language pretrained (VLP) representations with specialized local modeling, but it remains unclear whether the benefit of local modeling under strong global representations is tied to such specialized local modeling pipelines. Through controlled studies, we find that global VLP features provide a strong retrieval anchor and that, in a controlled ENCODER study, replacing its specialized local branch with lightweight token compression causes only a small degradation. On FashionIQ, dense and compressed local representations also perform similarly once global-local branch weighting is controlled. These observations motivate a narrower design question: how should a compact local branch allocate its limited capacity across reference evidence for each modification? We propose DART, a Distributed Information Bottleneck (DIB)-inspired module that conditions token-wise allocation on modification semantics and regularizes the allocation coefficients around a compression prior. DART is combined with balanced global-local aggregation while preserving the host VLP backbone and multimodal composition mechanism. Experiments on FashionIQ and CIRR support distributed local evidence allocation as a compact local modeling strategy under strong global representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.