From Solving to Judging: SpecialJudgeBench for Verdict Prediction and Executable Judge Repository Synthesis
Abstract
Competitive programming benchmarks typically evaluate programs generated by large language models (LLMs) using existing judges. Yet it remains unclear whether LLMs can perform the judging itself, especially for constructive, interactive, and communication problems whose correctness depends on constraints, protocols, and information isolation rather than output matching. To study this gap, we introduce SpecialJudgeBench, comprising 270 problems and 4,540 candidate programs, 903 of which have numerical reference scores. The benchmark assesses two complementary capabilities: predicting reference verdicts and scores for candidate programs, and synthesizing complete executable judge repositories from public problem specifications. We evaluate generated repositories on hidden candidate programs through binary verdict agreement, alongside build, protocol, and runtime checks. Across seven models, even the best verdict accuracy stays below 90% for each problem type; all models have higher false-rejection than false-acceptance rates, and their valid score predictions underestimate reference scores on average. Across eight model–scaffold configurations, execution completion reaches 90.9%–98.8%, but only 42.2%–55.9% of repositories correctly judge every hidden candidate in their evaluation pools. Whether generative models directly judge candidate programs or are tasked with building reliable, complex judging systems, achieving trustworthy evaluation remains a substantial challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.