acceptodds
Under review as a conference paper at ICLR 2027

Beyond Model-Internal Signals: Steering VLLM Attention with Predicted Human Saliency

Abstract

To reduce hallucinations in VLLMs, training-free methods steer attention toward relevant image regions, identified from model-internal statistics alone. In this work, we ask whether an external reference, human visual attention, can serve this role. To answer this question, we introduce and release VLFeedback-ETC, an eye-tracking corpus of 21 participants annotating preferences on VQA samples from VLFeedback, a widely used VLLM alignment dataset. Since gaze is unavailable at inference, we also introduce QaSalFormer, a question-aware saliency model for natural images, and benchmark it against existing predictors on our corpus. Using the corpus to compare human gaze with LLaVA-1.5 attention, we find that they are weakly correlated, with large differences between layers. We then propose GazeSteer, which adds predicted saliency to the pre-softmax attention scores over image tokens. On CHAIR, predicted saliency reduces object hallucination beyond a spatially uniform map that adds the same attention mass, and the reduction is not explained by shorter captions. Interestingly, the utility of a map is not determined by its global alignment with human gaze. Instead, useful maps combine image-specific structure with enough coverage. Moreover, predicted saliency does not replace cross-head consensus, but combining both is at least as effective as consensus alone, suggesting that external saliency is a complementary signal to model-internal ones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.