FineRich: Towards Holistic Fine-Grained Text-Rich Image Understanding via Versatile HTML-Based Synthesis
Abstract
Fine-grained Text-rich Image Understanding (Fine-grained TIU) is challenging for multimodal large language models (MLLMs) to accurately perceive and understand textual details from their long visual contexts. However, existing benchmarks remain limited in scope and cannot comprehensively evaluate these capabilities. In this paper, we introduce FRBench, a holistic benchmark that systematically covers all 90 combinations of six image types, three granularity levels, and five task domains, enabling a thorough evaluation for fine-grained TIU. To construct high-quality data, we develop a versatile HTML-based synthesis pipeline that renders diverse text-rich images and derives source-traceable TIU annotations. Specifically, we convert source images into HTML and edit evidence region related to the query to increase local information density. Using this pipeline, we further automatically build a 34K training set and propose FineRich, a novel two-stage training method for fine-grained TIU, which includes: 1) a region-to-image OPSD stage that first strengthens basic fine-grained text recognition; and 2) a multi-task GRPO training stage further specializes the model across diverse TIU tasks. Experiments show that FineRich achieves leading performance on FRBench while improving performance on other fine-grained TIU and general OCR benchmarks, demonstrating strong fine-grained capabilities and remarkable generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.