acceptodds
Under review as a conference paper at ICLR 2027

SysLoopBench: Evaluating Long-Horizon System Evolution by Autonomous Agents

Abstract

Autonomous agents hold promise for accelerating the development of scientific and engineering systems. However, existing benchmarks provide limited insight into whether agents can independently recreate the capabilities accumulated in expert-developed systems and advance beyond them. To address this gap, we introduce Sysloopbench, a benchmark that challenges agents to recreate and surpass expert-developed capabilities through long-horizon evolution of basic starting systems, retaining the original objectives and interfaces while independently developing their own algorithms and designs. Sysloopbench comprises 30 carefully designed tasks across mathematics, physics, biology, chemistry, engineering, and computer science. We construct each task from an expert-developed real-world system, extracting target capabilities, preparing a basic starting system and development environments, and establishing executable references for evaluation. Agents iteratively test and revise their systems using execution feedback, and we evaluate the resulting systems against expert-developed references on held-out workloads. Our evaluation reveals both the potential of long-horizon evolution to advance systems beyond expert references and the challenges of turning iterative improvements into lasting, generalizable gains across performance and efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.