Second-Order Response Laws for LLM Judges: From Prompt Stability to Decision Quality
Abstract
When an LLM judge changes its verdict after a prompt is reworded, the change may come from the wording or the randomness of another call. Confusing these sources can mislead model comparisons and response selection, yet separating them with few repeated calls is difficult. Plug-in prompt-dispersion scores combine both sources. We develop a framework that connects measurements of judge variability to decision quality. Second-order response laws describe how verdict distributions vary across prompts; crossing prompts with candidate orders separates prompt, position, interaction, and call variation. Agreement between distinct calls removes finite-call bias. We derive the prompt estimator's exact variance and establish matching upper and lower bounds for detecting weak prompt effects under a fixed, balanced design. Identical-input controls and an independent reference with more repeated calls validate the correction. On a matched panel, correction reverses the Qwen–Gemma-3 ranking by prompt sensitivity. We analyze over one million recorded categorical judgments across six judge models and three benchmarks. Cooling makes Qwen's repeated calls more consistent while its corrected prompt estimate rises nearly fivefold, from .011 to .054; the plug-in estimate rounds to .05456 at both temperatures. To examine what greater repeatability means for decisions, we evaluate independent groups of votes balanced across candidate orders. In Qwen, cooling produces more ties; changing the tie rule on the same votes shifts the estimated accuracy effect of cooling by 9.84 percentage points. An analytical model of position dependence explains how balanced votes can concentrate near a tie boundary and lower accuracy when ties lead to abstention, even when the judge favors the correct answer on average. The value of consistency therefore depends on alignment with correct answers, how votes are combined, and the cost of abstention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.