ExamInkBench: Benchmarking End-to-End Line-Level Student Answer Reconstruction from Complete Exam Pages
Abstract
Reliable automated assessment requires reconstructing students' complete responses from full examination pages rather than transcribing isolated handwriting. Existing benchmarks typically evaluate handwriting recognition, document parsing, and answer grounding separately, making coupled failure modes difficult to measure. Therefore, we curate a gold-standard corpus of 4,017 fully annotated samples from approximately 31,769 real mathematics examination papers, with 141 public samples forming ExamInkBench and 3,876 reserved for internal development and self-evolution. Each response is represented as an ordered sequence of answer lines grouped by question item. We design a two-level evaluation protocol that separately measures line-level extraction and ordering, as well as intra-line transcription fidelity. Evaluations of mainstream multimodal models reveal frequent answer-line omissions and incorrect question associations, while accurate local recognition does not guarantee correct page-level reconstruction. Motivated by these structured failures, we propose ExamInk-Evo, a failure-driven skill self-evolution method that mines recurrent errors, induces targeted reconstruction and verification skills, and retains updates only when held-out validation performance improves. Experiments with tool-using agents show that the method consistently improves both extraction and transcription, with larger gains under stricter line-matching thresholds. The evolved skills also transfer positively across agents, indicating that they capture reusable reconstruction procedures rather than agent-specific behaviors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.