Visual Context Composition: Rendering and Grouping Evidence for Multi-Image VLMs
Abstract
Multi-image tasks require vision-language models (VLMs) to relate images, labels, and instructions. Text can enter the language decoder directly or pass through the visual encoder when rendered in an image. Placing multiple elements in one image also allows them to be visually encoded together. We introduce Visual Context Composition (VCC), a training-free method that controls these processing routes while keeping images with their associated labels. Across five VLMs and four benchmark families, the selected VCC input configurations improve average task scores over baseline inputs. Giving each labeled example its own image performs best on average for fine-grained classification; separating task roles, such as instructions, references, and queries, performs best in the other benchmark families. Fine-grained gains persist for two models when individual image sizes or total input pixel counts are matched in separate comparisons. With rendered content held fixed, changing which task roles share an image also changes accuracy. We hypothesize that joint visual encoding represents task-relevant relationships. Assigned labels can be read from reference-image features. Token-removal experiments further show that information propagated from a region to other visual tokens contributes to predictions even after that region's own tokens are removed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.