Mixture of Graders: Adaptive Strategy Routing for LLM-Based Mathematical Evaluation
Abstract
LLM judges for mathematical reasoning are sensitive to the grading protocol they use. We compare 13 protocols, spanning direct scoring, structured chain-of-thought, and multi-call verification or aggregation, on two new expert-annotated competition-mathematics benchmarks with disjoint problem sources and candidate models: Putnam-AXIOM-Grading (540 problems; 1,080 candidate solutions) and OPC-Grading (500 problems; 1,000 candidate solutions). Across four open-weight judges, no single protocol is best everywhere, and the best choice can flip between benchmarks, a pattern that holds even under controlled prompt paraphrases and Holm-corrected pairwise tests. Building on this, we introduce Mixture of Graders (MoG), a judge-specific router that uses only the problem and reference solution to predict a distribution over grading protocols and combine their scores adaptively for each instance. MoG outperforms the strongest fixed protocol for every open-weight judge on both benchmarks, reducing relative MSE by 6.7 to 27.8% on Putnam-AXIOM-Grading and 3.0 to 13.1% on OPC-Grading. These results show that, for reference-based grading of LLM-generated competition-mathematics solutions, the most useful protocol depends on the instance, and learned routing can exploit this. Code is available at https://anonymous.4open.science/r/mixture-of-graders-01CC/README.md.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.