How VLMs Plan What's Next
Abstract
When a vision-language model describes an image, it names one object at a time and it rarely repeats itself, yet how the model plans this order is not obvious. We characterize the mechanism a VLM uses to decide where to look before it looks. The mechanism has four main parts. First, binding heads match the words of the question against the image and turn them into pointers to every matching image region. Second, the model binds each written answer to the region it came from, and turns it into a progress signal: a cursor, the position in the image up to which the list has been written. Third, a select operation chooses the nearest unwritten pointer after the cursor, or stops generating. Finally, gaze heads read the image at the chosen region and return a language-free concept, which language heads and the final MLPs turn into a word in the language of the question. We validate each part through causal patching in Qwen3-VL models on a synthetic visual retrieval task, pairing donor and receiver runs so that every alternative explanation predicts a different answer. We further show that the mechanism holds across model sizes, input formats and languages.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.