From Clues to Closure: Benchmarking Ultra-Long-Horizon Autonomous Computer Use through Evidence-Driven Gameplay
Abstract
As large language models gain strong GUI skills, computer use is becoming easy for them: recent models such as GPT-6 Astra and Claude Opus 5.5 saturate short-horizon GUI tasks. What still stands between them and real work is autonomy over long horizons, in tasks that set a goal but give no path and run long enough that the agent must maintain its own task state. We show that evidence-driven investigation can be used to build environments that demand this autonomy, and we state seven requirements on a task's evidence that make memory and multi-path exploration unavoidable. We introduce Clue2Closure, an ultra-long-horizon computer-use benchmark of three investigation games about one family: the commercial detective game The Roottrees are Dead, its expansion, and Scattered Roots, a new game that our pipeline builds to the seven requirements by splitting every scored fact into jointly required fragments and hiding them among thousands of documents. Six frontier agents score between 26% and 100% on the original games, where GPT-6 Astra fills all 333 blanks in 2,207 steps, but Scattered Roots remains unsolved: within 3,000 steps, no agent holds more than 19 of its 312 points. Their trajectories show that agents fail to act on what they already know, revise errors by trial and error, and get lost among open paths.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.