acceptodds
Under review as a conference paper at ICLR 2027

The Emperor's New Scaffold: Semantically Equivalent Harnesses Reshuffle the Leaderboard

Abstract

Every new language model arrives with a leaderboard that crowns it, and the community reads these tables as measurements of the models themselves. But does a leaderboard for agentic tasks measure the model, or the model together with the harness of prompts and protocol it was run under? We show that a one-word, semantically equivalent edit to the harness drops a flagship model's accuracy from 93.3 to 14.0 on code generation, below all five sub-2B models we tested. To test whether this is an isolated failure, we evaluate 24 models under 32 equivalent harness variants on three multi-turn agentic tasks (text-to-SQL, math, code). Nearly every pair of closely ranked models can be flipped by some equivalent rewrite, with swings far beyond resampling noise, and equivalent harnesses disagree on the top-3 in 47% to 84% of configurations across the three tasks. This instability lands exactly where it costs the most. Nobody consults a leaderboard to learn that a frontier model beats a sub-1B one; they consult it to choose between close competitors. Tracing every deviation through an exact decomposition, we find three failure modes, Delivery, Format, and Content; which one fails depends on which model meets which rewrite, and no harness in our grid treats all models impartially. Harness rewrites cannot move mountains, but they reshuffle the close races. We therefore release ScaffoldSpread, an eight-configuration audit that reproduces full-grid sensitivity at ρ=0.981, together with the parameterized harness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.