Align-R: Rethinking Residuals in Open-Vocabulary Semantic Segmentation
Abstract
In open‑vocabulary semantic segmentation (OVSS), the community commonly discards the final residual connections to bridge the adaptation gap between pretrained vision-language models and dense prediction, yielding better segmentation mIoU. We rethink this widely adopted heuristic and uncover an overlooked trade‑off: improved segmentation performance comes at the cost of recognition recall for categories present in the image. Pretrained residuals therefore carry meaningful pre-learned signals rather than pure noise, but align poorly with dense prediction objectives. From this observation, we identify two core failure modes: , where semantic responses are displaced from target regions; and , where residual features produce false activations for categories absent from the image. To address both misalignments, we present , which preserves residual information via instead of naive removal. Align-R performs spatial alignment through semantic-center-guided random exploration to recover fine-grained spatial cues, and concept alignment via reliable-anchor distillation to preserve correct category concepts and intra-class relationships. Evaluated across 8 OVSS benchmarks, Align-R yields consistent gains for both training-based and training‑free pipelines. Our results show that Align-R effectively mitigates residual misalignment and achieves a more favorable trade-off between recognition recall and segmentation performance. The code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.