acceptodds
Under review as a conference paper at ICLR 2027

Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence

Abstract

LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under reprompting, challenge, or sustained pushback. We introduce the Wiggle Framework, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and politicalresponse evaluation, emphasizing difficult or class-balanced subsets. At turn 10, task-average wiggle rates range from 19–71% under repeated static consensus pressure and 40–73% for a single adaptive LLM persuader; coverage under at least one of three persuaders is 62–91%. Successful flips are more often corrupting than corrective in aggregate, but a base-rate-controlled analysis shows that initially incorrect binary verdicts are conditionally more likely to flip than initially correct ones. Beyond the framework itself, we find that baseline jury majority strength is associated with item instability, although its predictive strength falls under held-out-judge evaluation. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.