acceptodds
Under review as a conference paper at ICLR 2027

More Reasoning Does Not Always Make a Better Judge: Understanding Inference-Time Compute in LLM Evaluation

Abstract

LLM judges are increasingly used to evaluate outputs requiring complex reasoning, yet it remains unclear whether additional inference-time reasoning improves their decisions or merely produces longer rationales. We investigate this question using a simple forced-continuation intervention as a diagnostic probe across open-domain, safety, mathematics, and code preference benchmarks. The strong open-source judges we test generally fail to benefit: additional reasoning often causes semantic drift and overconfident errors. To isolate the training conditions that enable useful scaling, we train a 7B judge with reflection-augmented supervised fine-tuning followed by reinforcement learning with verifiable outcome rewards, and analyze checkpoints throughout training. Across four benchmarks, this judge improves aggregate accuracy by roughly five percentage points over comparable open-source judges and achieves a 4.1% average relative gain from additional reasoning, whereas the baselines exhibit negative scaling. Checkpoint and ablation studies show that this capability emerges primarily during reinforcement learning rather than from reflective supervision alone. We characterize successful and failed scaling through semantic alignment to oracle rationales, entropy dynamics, and calibration. Successful scaling preserves semantic alignment and uncertainty during deliberation; failed scaling exhibits drift and premature entropy collapse. These differences yield substantially better calibration and support statistically valid risk-controlled selective prediction. Our results show that more reasoning does not inherently produce better judgments: converting extra compute into reliable evaluation is a learned property.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.