Same Results, Different Rankings: Measuring Evaluation-View Dependence in LLM Agent Benchmarks
Abstract
A benchmark can select different winners without changing any saved result. In DeepSWE, two configurations exchange first and twelfth place when the evaluation changes from single-attempt success to four-attempt task coverage. Such changes raise a measurement question: how much of a leaderboard depends on the evaluation view, and where does that dependence occur? We introduce the Structured Leaderboard Ambiguity Index (SLAI), which combines changes in model-pair relations with their ranking position and relative rank-gap amplitude across a declared family of views. Each contribution retains evidence from an actual pair of views. ParetoSweep computes the index exactly and returns these witnesses. Two theoretical controls establish information that selected conventional summaries omit: identical ranking summaries can accompany an SLAI ratio of \(n-2\), and duplicating an existing view leaves the entire ambiguity surface unchanged. Across ten saved evaluation panels, the full family can reveal substantially more weighted change than any fixed pair: BFCL yields 27.93% versus 9.36%. Four additional source panels retain both low-change cases and further family-coverage gaps. SLAI describes where view-dependent comparisons occur and preserves the evaluation conditions needed to inspect them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.