acceptodds
Under review as a conference paper at ICLR 2027

Equally Accurate, Unequally Legible: Measuring the Legibility of LLM Adjudicators

Abstract

Large language models are increasingly used to resolve real disputes, from e-commerce claims to prediction-market arbitration. An adjudicator is chosen when a contract is signed, long before the disputes it will rule on exist. At that moment, the only evidence about a judge is its record, usually summarized by reference agreement: how often its rulings match those of the institution it would replace (e.g., a human jury). This tells a party how closely the judge tracks that institution, but not whether the party can anticipate its rulings. Following the legal-realist view of law as the prediction of what judges will do, we introduce adjudicative legibility: how well an external forecaster can predict an adjudicator's ruling on an unseen dispute, given the dispute and, optionally, its published rules and past rulings. We study eight LLM adjudicators on contested e-commerce and article-deletion disputes. On more than a third of them, the outcome depends on which adjudicator is chosen, and adjudicators of similar accuracy differ by up to in illegibility. We prove that a party can protect itself by demanding a safety margin that grows with a judge's illegibility; on these disputes, this margin is comparable to that with no forecast at all. A party cannot meaningfully consent to a judge it cannot anticipate; we therefore argue that as LLMs take on binding adjudicative roles, legibility should be measured as routinely as accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.