Instances of equal computed difficulty differ in LLM agent success by more than chance
Abstract
Agent benchmarks usually set difficulty by how many models failed a task, building the axis from the systems measured: a fall in success along it is not itself evidence about a model. Does difficulty computed from the task's own structure predict how often an LLM agent wins? An instance is one setup in a small autobattler: the agent buys and places three units, one at a time, against a script. Its 3,240 complete plays can all be listed, and its difficulty d is the fraction that win: larger d means an easier instance. Every unit costs the same and the agent spends its whole budget, so choosing at random at each step draws uniformly from those plays: random play wins with probability exactly d. Two LLM agents, luna and glm, played 126 instances from the range's two ends, as did a third model reported separately, a random player and 21 scripted strategies. An analysis plan fixed four Holm-adjusted tests in advance; all four reject their null hypotheses. Success rises with computed d (p = 0.0004), and luna and glm each beat random play's 0.100 at d = 0.1, one test each. No instance sits at d = 0.1, fixed in advance between the ends, so both come off a fitted curve. The fourth rejects too: instances of equal computed difficulty differ widely in success, beyond what difficulty explains; structure is not all that makes an instance hard for a model. Instance for instance, luna and glm respond to difficulty about as steeply as random play (90% interval 0.92 to 1.30, within the pre-registered range 0.6 to 1.4). That steepness is a slope, 1 for random play by construction. The plan's model assumed one shared slope; the agents' separate fits differ. The slope does not tell you how well a strategy plays: four scripted strategies spanning the design's range of information and search have slopes within about 0.05 of one another and of random play's fitted 1.03. Computed difficulty does predict success, but is no substitute for measuring models: a difficulty-graded suite should report the spread at equal difficulty beside the slope. These results rest on two agents, one game and one prompt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.