Look Closer, Segment Better: Query-Guided Visual Token Unfolding
Abstract
Multimodal large language models (MLLMs) demonstrate strong vision–language understanding capabilities, yet precise pixel-level segmentation remains challenging. Existing methods primarily focus on mask representation and generation, while the visual tokens fed into the large language model are typically spatially compressed, weakening local details and limiting segmentation accuracy for complex objects. To address this issue, we propose query-guided ***V***isual ***T***oken ***U***nfolding (***VTU***), which enables fine-grained visual tokens to directly perform in vision–language interaction before mask prediction. ***VTU*** first processes coarse-grained visual tokens to achieve a global understanding of the image and the text query, and employs a lightweight router to select tokens for unfolding. The corresponding tokens are inserted before the segmentation token, allowing the segmentation representation to integrate global context and local details. We further introduce router supervision targeting mixed-region and a simple-sample skipping strategy to incorporate informative tokens while reducing unnecessary unfolding. Extensive experiments demonstrate that ***VTU*** achieves state-of-the-art performance across multiple segmentation tasks and different model scales with only a small number of additional visual tokens, while preserving general visual understanding capabilities. Anonymous link: [https://anonymous.4open.science/r/VTU](https://anonymous.4open.science/r/VTU).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.