acceptodds
Under review as a conference paper at ICLR 2027

ExamOCRBench: A Multi-Task Benchmark for Real-World Exam Document Understanding and Automated Grading

Abstract

Existing evaluations of vision-language models cover text recognition, document parsing, and visual question answering, yet end-to-end grading of student-completed examination pages remains underexplored. This setting requires faithful interpretation of handwritten answers and corrections in the context of questions and page layouts to support grading. We introduce ExamOCRBench, built from photographed primary-school examination pages in Chinese, Mathematics, and English, with linked structural, transcription, and human-verified grading annotations. It defines four tasks: exam paper grading, structure parsing, content grounding, and handwritten text recognition. End-to-end grading forms the core evaluation, while the remaining tasks independently diagnose related capabilities. Across five evaluated models, complete-page grading accuracy ranges from 33.22% to only 58.18%, highlighting the difficulty of consistently assessing all student responses on a page. Structure recovery, content grounding, and handwriting recognition also remain challenging. Qualitative analysis also reveals that misreading an incorrect student response as the expected answer can lead to undeserved credit. ExamOCRBench provides a unified framework for evaluating real-world examination grading and identifying related capability gaps.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.