Detecting Is Not Mitigating: A Benchmark for the Awareness–Behavior Gap of LLM Judges under Score Manipulation
Abstract
LLM-as-judge systems are increasingly deployed to score model outputs, yet they can be manipulated by content embedded in the answer being graded (injected grading instructions, self-praise, fabricated authority or consensus, unfounded confidence, rubric hijacking, padding). Prior work measures whether such manipulation raises scores, but conflates two distinct capabilities: whether a judge detects manipulation and whether its score is immune to it. We introduce a benchmark that decouples these two axes with a dual-probe protocol, and a single diagnostic metric, the D-gap, defined as the net rate at which a judge detects manipulation yet still inflates the score. Across 6 frontier judges and 4000 items spanning math, code, and factual domains (all with confirmed-wrong answers, so the ideal D-gap is zero), we find detection is near-perfect (rate ≈1.00) yet D-gap is significantly positive for every judge, revealing a systematic awareness–behavior gap: judges know they are being manipulated but grade as if they do not. We further show strong domain heterogeneity and, as a first partial solution, that an explicit anti-manipulation instruction cuts mean D-gap by 27%—nearly eliminating it for the strongest judge on injection—while remaining stubbornly high for others, establishing the gap as a real yet only partially solved challenge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.