acceptodds
Under review as a conference paper at ICLR 2027

Scene-Graph Molmo: Adding Structure to Pretraining Yields Strong Vision-Language Models

Abstract

Pretraining on large, diverse collections of image-caption pairs is the standard method for training strong Vision-Language Models (VLMs). However, captions have several drawbacks due to their lack of consistent structure: they can be brief, ambiguous, and leave out information that annotators think is “obvious”, leading to failures on downstream grounding and reasoning tasks. Motivated by human visual processing, we instead use a structured scene description: scene graphs. We curate PixMo-SceneGraph, a dataset of 704K image-scene-graph pairs derived solely from captions in the PixMo-Cap dataset, enabling a controlled comparison with unstructured caption data. We pretrain Molmo2 on PixMo-SceneGraph and finetune the resulting models on 8 standard academic datasets, finding that scene- graph pretraining improves performance over caption pretraining, particularly in generalizing to the grounding and reasoning tasks where caption-based models struggle. Combining scene graphs with captions yields further gains. Moreover, the structure of scene graphs also provides an interface for targeted data augmentation: we augment PixMo-SceneGraph with information typically omitted from captions—temporal statements, negations, affordances, and external knowledge—further improving performance. Our suite of models, Molmo-SceneGraph, outperform Qwen3-VL at 2B, 4B, and 8B scales across a broad range of academic, grounding, and reasoning benchmarks, despite being trained on 200–500× fewer tokens. Our results highlight the value of structure in pretraining data, and encourage researchers to rethink standard methods for VLM pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.