Mask–Language Cycle Consistency for Test-Time Adaptation in Reasoning Segmentation
Abstract
Reasoning segmentation aims to predict pixel-level masks from complex, open-ended language queries. Despite recent progress, vision–language models remain vulnerable to test-time distribution shifts, such as environmental changes and image corruptions, which can disrupt mask–language alignment and degrade segmentation quality. Existing test-time adaptation (TTA) methods for segmentation primarily rely on entropy minimization or self-training, treating segmentation as a one-way prediction problem while overlooking its bidirectional cross-modal structure. We propose Cycle-ML, a TTA framework based on mask–language cycle consistency. At inference time, Cycle-ML predicts a mask from the input query, generates a referring expression grounded in the predicted mask, and regrounds the generated expression to obtain a cycle-reconstructed mask. It then uses both mask-level and language-level consistency as self-supervised adaptation signals. To mitigate error amplification from incorrect yet self-consistent predictions, we further introduce consistency–reliability learning, which estimates sample reliability from joint mask and language consistency and adaptively weights the resulting supervision. Experiments on multiple reasoning and referring segmentation benchmarks under 15 synthetic corruptions show that Cycle-ML consistently outperforms strong TTA baselines, demonstrating the effectiveness of reliability-aware mask–language cycle consistency for robust vision–language grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.