Position, Not Content: What Repairs Question-First Prompting in Vision-Language Models
Abstract
Putting the question before the image rather than after it changes nothing about what a vision-language model is told, yet it changes what the model answers. Prior work by the present authors established this question-first paradox on three yes/no benchmarks and traced it to read-out rather than perception: the question does reach the image and visibly steers it, but the answer position cannot read it back, and repeating either the question or the image repairs the damage. What that work could not say is how far the effect reaches, what the repeated image actually contributes, and what it costs. We answer those three questions. Across 14 further benchmarks the paradox proves general, costing 7 points on average and as much as 18 on a single benchmark, yet it is not a law: on object detection, where the model returns boxes instead of a word, the effect reverses and asking first is the better choice. The repeated image turns out to carry almost no information. Shrinking it to 1 8 of its resolution leaves accuracy statistically unchanged while cutting its price from thousands of extra tokens to about 100, and a reversed copy works as well as a correct one, so what matters is where the tokens sit and not what they show. Breadth also marks a limit: on video, repeating the image stops helping, while repeating the question still does. A reflective prompt optimizer given the same budget improves only the ordering our diagnosis says has room left, and never catches the simpler fix, so this is not something better wording would have found. The practical consequence is immediate: put the question after the image, and where it has to come first, repeat it afterwards. The consequence for the field is that prompt layout, which papers rarely report, moves benchmarks further than the model differences those benchmarks are built to measure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.