ImageParser: Hierarchical, Grounded Image Parsing with Small Vision–Language Models
Abstract
We investigate image perception as reconstruction in language: preserving the descriptive information needed to reconstruct an image up to semantic equivalence. ImageParser, a family of 0.8B, 2B, and 4B vision–language models, converts images into hierarchical, spatially grounded JSON parses. Each parse links global content and appearance to localized entities and parts, preserving visible text and explicit parent relationships without predefined categories or downstream questions. We train ImageParser on 6.1 million image–parse pairs produced by frontier perception models and iterative multimodal verification, averaging over 3,000 annotation tokens per image. On ImageParserBench, a human-reviewed benchmark of 1,975 images, ImageParser-4B leads all evaluated general-purpose VLMs, including Gemini 3.8 Flash and GPT-5.6-sol, across six parsing metrics. It improves caption and hierarchy F1 over their strongest external baselines by 20.0 and 11.8 percentage points, respectively. Even the 0.8B model surpasses the evaluated proprietary models in semantics, type, and hierarchy. Task-specific readouts reuse these parses for detection, OCR, and captioning; reconstruction and pixel-free QA assess information preservation and remaining limitations. These results show that compact VLMs learn reusable hierarchical image representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.