acceptodds
Under review as a conference paper at ICLR 2027

ImageParser: Hierarchical, Grounded Image Parsing with Small Vision–Language Models

Abstract

We investigate image perception as reconstruction in language: preserving the descriptive information needed to reconstruct an image up to semantic equivalence. ImageParser, a family of 0.8B, 2B, and 4B vision–language models, converts images into hierarchical, spatially grounded JSON parses. Each parse links global content and appearance to localized entities and parts, preserving visible text and explicit parent relationships without predefined categories or downstream questions. We train ImageParser on 6.1 million image–parse pairs produced by frontier perception models and iterative multimodal verification, averaging over 3,000 annotation tokens per image. On ImageParserBench, a human-reviewed benchmark of 1,975 images, ImageParser-4B leads all evaluated general-purpose VLMs, including Gemini 3.8 Flash and GPT-5.6-sol, across six parsing metrics. It improves caption and hierarchy F1 over their strongest external baselines by 20.0 and 11.8 percentage points, respectively. Even the 0.8B model surpasses the evaluated proprietary models in semantics, type, and hierarchy. Task-specific readouts reuse these parses for detection, OCR, and captioning; reconstruction and pixel-free QA assess information preservation and remaining limitations. These results show that compact VLMs learn reusable hierarchical image representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.