acceptodds
Under review as a conference paper at ICLR 2027

Prompt Lottery in LLM-As-a-Judge: Measuring and Mitigating Sensitivity to Syntactic Rephrasing

Abstract

LLM-as-a-Judge (LLMaJ) has become a standard paradigm for evaluating large language models, yet most approaches implicitly assume that a single judge prompt functions as a stable evaluation instrument. Prior work has primarily explored deliberate prompt modifications such as scoring rubrics, role descriptions, or chain-of-thought reasoning. In contrast, we investigate whether validated prompt variants intended to preserve evaluation semantics while differing in surface realisation (paraphrastic rephrasing) yield consistent judgments under a fixed evaluation task. We introduce Prompt Lottery, a phenomenon in which prompt reformulations intended to preserve semantic content produce different LLMaJ judgments. Across a large-scale study spanning diverse evaluation tasks, judge families, and model backbones, we find that paraphrastic variation can alter evaluation outcomes and, in some settings, reverse relative model rankings. We conducted a user study and show that variability of LLMaJ exceeds the observed variability in out study, as well as generally accepted variability in human subject studies raising concerns about the stability of single-prompt evaluation. Finally, we show that multi-prompt aggregation substantially improves robustness by emphasising judgments that remain consistent across prompt variants.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.