The Honest Agent Paradox: How LLM Judges Score a Human-Inspired Web Agent
Abstract
LLM judges decide most published results for web agents and are usually validated by test-retest self-consistency. We show that self-consistency cannot detect a large, systematic error, the Honest Agent Paradox: when the same verified result is reported in different words, LLM judges change their verdict, and the direction of the change depends on the judge and its prompt. We split it into Strictness (α), the false-negative rate on correct results reported with disclosed execution friction, and Gullibility (β), the false-positive rate on incorrect results reported as complete, which can offset in aggregate success rate while individual verdicts stay wrong. In a controlled probe of constructed reports whose correctness is known by construction, disclosing friction on 65 correct reports lowers the pass rate of gpt-4o, claude-3-5-sonnet and gemini-1.5-pro by 38.5 to 46.2 percentage points (p < 1e-5), and hedged wording alone costs 13.9 to 23.1 points; on 21 incorrect reports, confident assertion turns up to 38.1% of failures into passes. The production WebJudge moves the other way (disclosure 65/65, assertion 56/65, exact McNemar p = 0.00391), so rankings produced by different judges cannot be compared. We describe the agent harness behind our memory study, whose vision grounding, bounding-box snapping and trusted click cascade operate the page, and whose recovery machinery, which reports partial results with caveats, is where such disclosures originate. In a 240-run memory study on Online-Mind2Web, double-blind human adjudication shrinks an apparent 25.0 point gap between two memory architectures (11/24 vs. 17/24) to 8.3 points (18/24 vs. 20/24), and 10 of the 12 discordant tasks trace to how the reports were styled. We release the framing probe at https://anonymous.4open.science/r/honest-agent-paradox.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.