acceptodds
Under review as a conference paper at ICLR 2027

FoAReD: Foreground-guided Attention Reallocation Decoding for Comprehensive Image Captioning

Abstract

Despite their strong image captioning capabilities, VLMs often fail to capture fine-grained details in complex scenes. We find that this issue is associated with insufficient utilization of visual information during autoregressive decoding, where attention is dominated by redundant background and sink tokens. To address this issue, we propose FoAReD, a training-free image captioning framework that enhances visual information utilization through attention contrastive decoding. Specifically, FoAReD employs foreground-guided attention reallocation to construct a contrastive attention condition. During autoregressive decoding, the contrastive and original attention conditions are contrasted to enhance the generation preference toward informative visual details. Experiments on CompreCap and CaptionQA across multiple VLMs show that FoAReD consistently improves both caption comprehensiveness and accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.