acceptodds
Under review as a conference paper at ICLR 2027

LLM-as-a-Verifier: A General-Purpose Verification Framework

Abstract

Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of large language models (LLMs). In this work, we identify verification—the ability to determine the correctness of a solution—as a new scaling axis. To unlock this, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agents without requiring additional training, and systematically characterize how verification scales. Specifically, we identify three orthogonal dimensions of verification scaling: (1) score granularity, (2) repeated evaluation, and (3) criteria decomposition, and show that they provide complementary gains and should be scaled in tandem to achieve optimal performance. In particular, we show for the first time that scaling score granularity improves the separation between correct and incorrect solutions, yielding better-calibrated comparisons than the conventional 1–5 scale. To make verification scaling practical, we propose Probabilistic Pivot Tournament, a compute-efficient ranking algorithm that mitigates positional bias and adaptively allocates verification compute toward stronger and uncertain candidates, reducing the number of pairwise comparisons from to . When used as a trajectory reward model for test-time scaling, LLM-as-a-Verifier achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%). We also find that pairing LLM-as-a-Verifier with an open-weight model outperforms closed-source frontier baselines on Terminal-Bench 2.1 with substantially lower cost. Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build an extension for Claude Code, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can be used as a dense reward signal for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks. Interestingly, our approach also improves the effectiveness of RL in non-verifiable domains without ground-truth checkers, such as painting, substantially outperforming a standard LLM-as-a-Judge baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.