acceptodds
Under review as a conference paper at ICLR 2027

Rewarding What Matters: Human-Inspired Visual Search for Fine-Grained MLLM

Abstract

Multimodal large language models (MLLMs) have made substantial progress in visual understanding, yet they still struggle to resolve small but answer-critical details in high-resolution images. Exhaustively processing those images can preserve such details but incurs substantial visual-token cost, motivating selective visual inspection. In contrast, human visual search efficiently prioritizes information through complementary cognitive mechanisms: bottom-up processes that suppress irrelevant distractions and top-down cognitive control that directs attention toward task-relevant information. Inspired by these mechanisms, we introduce SAGE (Surprise-And-Goal-Guided Exploration), a training-free visual search method with a bottom-up surprise module and a top-down relevance module. The surprise module identifies more informative visual tokens relative to their surrounding context, while the relevance module selects regions that are most semantically relevant to the task query. Together, these modules enable efficient filtering of redundant visual content while preserving task-critical evidence. Experiments across four MLLM families demonstrate that our method consistently improves accuracy over ZoomEye, while requiring fewer visual tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.