acceptodds
Under review as a conference paper at ICLR 2027

Stop Describing, Start Answering: Reducing Description Reliance in Vision-Language Models

Abstract

Existing Vision-Language Models (VLMs) have an undesirable habit in visual question answering: models often rely on image descriptions, first describing the visual content and then answering the question. This describe-then-answer pattern can improve accuracy, but it also produces redundant tokens and increases latency. When forced to answer directly, VLMs give shorter but often less accurate answers. We call this gap description reliance. To reduce this reliance without generating descriptions at inference, we propose Internalization, a two-stage training method that uses caption-derived answer supervision to train a compact image-question prior for direct answering. Stage I aligns the new prior module with the base model distribution through logits-wise warm-up, and Stage II injects prior tokens into hidden states through preference training, using the Stage I model as a teacher to reduce distribution drift. We further introduce InternVQA, a benchmark that measures description reliance across image settings, task types, and answer lengths. Experiments on Qwen3-VL and LLaVA-OneVision show that Internalization raises the average direct-answer score from to and shrinks the caption-direct gap from to judge points while increasing the caption score from to . It also achieves aggregate retention across ten public benchmarks and shortens end-to-end inference time by by generating fewer tokens, showing that caption-derived supervision can improve direct answering through compact prior tokens instead of generated descriptions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.