Dense Segmentation as First-Order Output of Vision-Language Models
Abstract
Autoregressive vision-language models (VLMs) exhibit fine-grained spatial reasoning when describing images in words. However, producing the corresponding segmentation masks usually requires external segmentation models or learned mask decoders. This work introduces FOMO (first-order mask output) a geometry language that makes dense segmentation a native autoregressive output. FOMO's deterministic encoder uses adaptive quadtrees and pre-defined geometry stencils to describe binary masks as token sequences. Homogeneous regions are compactly encoded with only a few tokens, while additional tokens are allocated to complex boundary regions. FOMO sequences generated by VLMs deterministically form a mask without neural decoders. Theoretical proofs and empirical experiments demonstrate the codec's near-exact reconstruction, a Spearman correlation of between token budget and mask quality, and graceful degradation under most grammar-preserving token corruptions. Moreover, with the FOMO vocabulary extension and fine-tuning via low-rank adaptation, Qwen3.5-9B learns to generate masks for referring-expression segmentation queries without requiring any changes to its vision encoder or added external modules. Our method achieves 76.0%, 71.7%, and 71.7% cumulative IoU for RefCOCO val, RefCOCO+ val, and RefCOCOg val, using fixed geometric reconstruction without an external segmentation model or learned mask decoder.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.