Can VLMs Assemble Local Decisions into Global Structures?
Abstract
Many vision-language tasks require combining multiple local decisions into a single structured output. We ask whether vision-language models (VLMs) can reliably assemble such decisions into assignments, orderings, groupings, and hierarchies. We identify an assembly deficit: a model can answer the underlying local questions accurately yet fail to produce the corresponding global structure. On RecipeQA, we isolate this deficit while holding the model and task input fixed, finding that local step-image decisions are substantially stronger than whole-order prediction or scoring. We introduce local+solve, a training-free procedure that extracts local scores from a frozen VLM and assembles them with a task-specific solver under task-defined constraints. Across six controlled structural families, five tasks drawn from public benchmarks, and GUI-agent settings, local+solve improves over matched global prediction, including stronger inference variants where tested, and its gain operationally measures the recoverable part of the deficit. The deficit narrows from 7B to 72B but persists in most families. Post-training on local+solve's outputs nearly internalizes assembly on the trained task, while multi-family local+solve pseudo-labels improve global prediction on two unseen families, suggesting that assembly is a distinct, learnable capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.