Boundary-Informed Support-Guided Evidence Fusion for Training-Free Multimodal One-Shot Segmentation
Abstract
How should category text and visual support be combined when they disagree about the identity or extent of a query object? We introduce Support-Guided Evidence Fusion (SGEF), which treats the annotated support as both a prompt and a reliability reference. A frozen SAM3 evaluates one query-text view and two layouts of the same one-shot support-query pair. Their aligned fields form a semantic-rescue branch and a layout-consistent signed-correction branch, while support-region overlap and inner-outer boundary fidelity determine an episode-specific mixing weight. This boundary-informed rule adds no learned parameters and requires no task-specific training or test-time optimization. Across eight benchmarks, SGEF reaches 89.28, 77.42, and 61.82 foreground mIoU on PASCAL, COCO, and LVIS, and obtains the highest score among the evaluated methods on four of five additional benchmarks. Component replacements and episode-level weight reassignment isolate the contributions of directional correction and support-boundary fidelity. Across six frozen encoder-decoder configurations evaluated on complete PASCAL and COCO fold 0, fusion improves the corresponding reference field in all twelve comparisons by 0.57-1.85 mIoU points. These results establish annotated support boundaries as an effective control signal for training-free multimodal one-shot segmentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.