acceptodds
Under review as a conference paper at ICLR 2027

A Tale of Two Signals: Open-vocabulary Dense Perception with Frozen MLLMs

Abstract

Multimodal large language models (MLLMs) encode spatial semantics in decoder hidden states and provide image-conditioned responses about category presence. We investigate how these signals complement each other for spatial recognition without task-specific fine-tuning. Our analyses show that same-layer visual–text pairing is not necessarily optimal, and selected visual-layer aggregation can improve recognition beyond the strongest single-layer counterpart in our diagnostic. We further distinguish the contextual role of jointly encoded categories from their role as prediction alternatives: textual context affects discrimination, while category competition can alter recognition even with fixed representations. Building on these observations, we introduce TALE (Semantic Aggregation and Selective Verification), which aggregates visual states, jointly constructs contextual category representations, and independently verifies leading candidates in ambiguous regions. Valid negative responses guide similarity refinement without reconstructing the classification representations or restricting the full prediction vocabulary. Experiments on regional classification, open-vocabulary detection, and segmentation demonstrate the utility of TALE, while controlled ablations characterize its recognition gains and inference costs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.