ReMarkBench: Benchmarking Agentic Grading of Raw Examination Documents
Abstract
Autonomous grading agents that assess student work with minimal human intervention can reduce teachers' workload and provide timely, informative feedback at scale. However, existing benchmarks mostly focus on pre-segmented student answers and check whether scores are correct, not whether each score rests on the student's answer and the rubric, or whether its grounds stay the same when the answer is regraded. To address these gaps, we introduce ReMarkBench, a benchmark for evaluating end-to-end agentic grading from raw examination documents: complete question papers, grading rubrics, and handwritten answer scripts from a public university-entrance examination. It covers 3,570 grading units in six grading tasks across Mathematics, Physics, Chemistry, Biology, and English. For every grading unit, the grading agent must produce a score and a supporting justification that cites the student's actual answer and explains the score under the rubric. We therefore assess three complementary dimensions: score accuracy, justification quality, and cross-run consistency. Across five frontier grading agents, each a model–harness configuration reading the raw documents directly, even the best reaches a joint pass rate of only 41.0%, averaged equally across the six tasks over three independent runs, and many correctly scored records still fail justification quality. Substantial gaps remain even for an agent grading from human-verified structured inputs instead of the raw documents, so the difficulty goes beyond reading the documents. ReMarkBench provides a testbed for diagnosing these gaps and advancing accurate, interpretable, and reproducible automated grading.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.