acceptodds
Under review as a conference paper at ICLR 2027

QuoteBench: When Matched Scores Hide Command-Path Failures

Abstract

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench separates the two by replay. It holds a model's reply fixed and executes it on two transports, direct and through one deliberately unescaped added parser, and crosses that with two generation contracts, one silent about the boundary and one that discloses it. The matched gap between the two paths then decomposes exactly, per task, into transport damage and contract-conditioned compensation. Across eight same-window configurations on 56 exact-final-state tasks from 14 incident-derived families, replaying the same reply through the added parser loses 55.4–73.2 points; disclosure recovers 30.4–60.7 points for six configurations and zero or slightly negative for the other two, and three public draws preserve both ranges. Raw generation is nearly saturated at the frontier, so boundary adaptation is what separates models: at the frontier the two effects cancel (GPT-5.6-sol's matched gap of −3.6 points hides −64.3 damage and +60.7 compensation), while elsewhere the matched score falls with the path, and the deployment configuration reorders models. A real ssh boundary reproduces the damage, and correct escaping or a temporary script removes it. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator instead of treating a matched score as an intrinsic model property.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.