acceptodds
Under review as a conference paper at ICLR 2027

GLASS: Unlocking Dense Perception in Generative Vision Encoders

Abstract

Generative vision encoders pretrained with an autoregressive caption objective transfer well to multimodal LLMs, but their patch-level representations lag on dense perception. Caption prediction does not explicitly supervise the geometry or absolute position of individual patches. In our transfer experiments, existing dense post-training recipes improve dense prediction but reduce downstream captioning and multimodal understanding. We explore retaining the caption objective while adding masked feature recovery. Here the mask selects both the recovery targets and the visual evidence available for caption prediction. Hiding the patches receiving the most caption attention produces a failure pattern we term text prior collapse. Features from different images become more similar and dense performance declines, yet the caption loss remains close to the other masking runs and does not reflect the extent of the degradation. We propose Grounded LAyout and Semantic maSking (GLASS), a caption-guided dense post-training recipe. Semantic masking keeps the highest-attention half of the patches visible, and masked feature recovery uses this context to predict full-image teacher features at hidden locations. Layout recovery adds explicit geometric supervision by re-encoding the visible patches without positional information and regressing their normalized coordinates. On GenLIP and OpenVisionĀ 2, GLASS improves dense perception and downstream captioning, with near-pretrained ImageNet accuracy and aggregate multimodal scores under the evaluated protocols. Further dense evaluations show gains across model scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.