acceptodds
Under review as a conference paper at ICLR 2027

Rubric Optimization for LLM-Based Automated Grading with Data Synthesis

Abstract

Large language models (LLMs) have recently emerged as a promising approach for automatic grading (AG) due to their strong instruction-following and reasoning abilities. Recent work has explored automatic rubric optimization using labeled student responses to address rubric quality issues. However, its effectiveness is often limited by the scarcity of annotations. In this work, we propose a new optimization paradigm, RODS, which leverages the generative capabilities of LLMs to synthesize responses as supplementary data, enabling existing rubric optimization methods to perform robustly under limited-label settings. To mitigate noise in synthetic labels, we design a consistency-based optimization mechanism. We further introduce a generative adversarial-based quality control (GAQC) module to improve the fidelity of synthetic samples, reducing distribution shift and enhancing alignment with real student responses. In addition, we integrate both synthetic and scarce human-labeled data within an iterative optimization framework, ensuring that the optimization process remains grounded in improving grading performance on real responses. Experiments on multiple AG datasets demonstrate that RODS consistently improves rubric quality and grading accuracy, highlighting its effectiveness and practical potential for real-world LLM-based automatic grading.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.