Focus where it Matters: Training-Free Semantic Re-Focus for Vision Language Models
Abstract
Vision-language models (VLMs) often fall back on linguistic priors even when the visual input contains the evidence needed to answer a query. We systematically investigate this failure through controlled input perturbations and decoder attention analysis. While visual context is essential for accurate generation, visual tokens receive disproportionately low attention. Moreover, this attention is concentrated in a small subset of heads and layers and is often poorly aligned with query-relevant image regions. Motivated by these insights, we introduce Semantic Re-Focus (SRF), an inference-time method that intervenes at two complementary stages. A semantic relevance map is estimated and is used to apply foveation before visual encoding and a selective, query-conditioned enhancement of visual attention is performed during decoding. Across five benchmark suites and two VLM architectures, SRF improves performance in fine-grained perception, counterfactual settings, object hallucination, and open-ended generation. On Qwen2.5-VL-3B and LLaVA-1.5-7B, it improves MMVP pair accuracy from to , and lowers the MMHal-Bench hallucination rate from to respectively. Our results suggest that query-conditioned control of visual evidence is a promising direction for building more reliably grounded multi-modal models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.