Text-Conditioned Visual Steering for Hallucination Mitigation
Abstract
Multimodal large language models (MLLMs) are becoming a general interface for visual question answering, captioning and image-grounded reasoning. Their fluency makes them useful in practice, while also making unsupported objects and relations difficult to detect. We evaluate our method on POPE, CHAIR and complementary multimodal benchmarks. Recent inference-time methods use contrastive decoding, attention intervention, visual-token reinjection or latent steering. A representative class of visual-steering methods contrasts image-conditioned and text-only activations, extracting a reusable direction for generation. Yet this direction typically compresses the whole visual sequence into one global shift, leaving the text-dependent organization of visual evidence implicit. As a result, a correction that is useful for one question can be poorly aligned with another question about the same image, while methods that preserve token-level structure often add decoding-time computation. To address this gap, we introduce (Text-Conditioned Visual Steering), which combines global image–text contrast with text-conditioned visual-token routing. TVS uses the attention weights from the current text query to estimate which visual tokens should receive more or less emphasis, reweights them around a uniform reference, and pools their hidden states with signed zero-sum coefficients. The resulting correction is added to the inherited global direction, so the method remains training-free, preserves the steering interface and caches one direction for ordinary decoding. Across POPE, CHAIR and general multimodal evaluations, TVS improves the visual-steering baseline on Qwen and on LLaVA MSCOCO, with smaller but positive average gains on LLaVA A-OKVQA and GQA and variation across individual splits. Component ablations verify that the global and token-level terms are complementary, while efficiency experiments show that the cached implementation retains reference decoding accuracy with lower overhead. Code is available at https://anonymous.4open.science/status/TVS-E8F0.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.