VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) struggle with fine-grained visual search, where a question concerns a small object in a high-resolution image. Existing remedies fall into two families. Training-free methods zoom in using attention or confidence scores but query the MLLM several times per example. Models trained with reinforcement learning (RL) to call cropping tools replace these queries with multi-turn rollouts but their tool calls are hard to control or interpret. To close this gap, we propose VisLens (Visual Focus via Logit Lens), which locates the target by reading the MLLM's own visual tokens through the logit lens, the projection of a hidden state through the LLM head onto the vocabulary. VisLens adds a lightweight tuned lens that maps early hidden states into the late hidden-state space, so visual tokens can be already read out from early layers. These tokens are matched to the target words in the query to produce a crop of the relevant region, which is fed back in alongside the original image to generate the final answer. The whole process, from decoding to answer, completes in a single forward pass with no repeated queries. Across three backbones, VisLens gains to points on VBench and HR-Bench and lies on the accuracy-latency Pareto frontier. It runs - faster than RL tool-use models and up to faster than training-free multi-pass search methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.