acceptodds
Under review as a conference paper at ICLR 2027

Dianjin-DocBench: A Full-Spectrum Benchmark for Financial Document OCR

Abstract

Financial document parsing must preserve both the content of individual pages and the continuity of information fragmented by layout and pagination. We introduce Dianjin-DocBench, a benchmark covering 712 documents and 1,033 pages across 16 financial and finance-related categories in Chinese and English. Its manual annotation workflow separates layout analysis, content recognition, and truncated-content linking, producing page-level ground truth with unmerged regions and document-level ground truth with explicit relations and correctly consolidated text and tables. The evaluation protocol covers text, formula, table, and image parsing, alongside separate same-page and cross-page continuation tasks. We propose a new image evaluation metric that measures overlap between unions of predicted and annotated regions, accommodating differences in region granularity while penalizing both missed visual content and excess predicted area. We evaluate pipeline-based specialized VLMs, end-to-end specialized VLMs, and general VLMs under this unified protocol. By separating recognition, visual preservation, and continuation, Dianjin-DocBench provides a focused evaluation of the information retained in parsed financial documents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.