Beyond Selection : Integrating Complementary Information for Training-Free Visual In-Context Segmentation
Abstract
Recent visual foundation models provide strong visual representations and segmentation capabilities, enabling training-free visual prompt-based in-context segmentation (VP-ICS) for novel segmentation tasks. However, existing VP-ICS methods predominantly rely on winner-take-all selection strategies, which retain only a limited subset of retrieved evidence and can discard complementary information distributed across multiple support shots. We propose Integra-ICS, a training-free VP-ICS framework that integrates support evidence to exploit such distributed information. For each query token, Integra-ICS retrieves the global top- support tokens from a unified support token pool containing both foreground and background. By applying masked softmax aggregation to query-support token similarities, Integra-ICS integrates complementary evidence across support shots. Candidate foreground regions are subsequently localized and filtered by a minimum-evidence criterion to obtain reliable predictions. Across eight diverse benchmarks, Integra-ICS outperforms the strongest baseline by 3.60 and 3.56 mIoU points on average in the 5-shot and 10-shot settings, respectively, using a single hyperparameter configuration without dataset-specific tuning. The improvement comes from integrating complementary cross-shot evidence, while the minimum-evidence criterion retains reliable predictions supported by foreground tokens from the support set.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.