acceptodds
Under review as a conference paper at ICLR 2027

GeneralsArena: Evaluating End-to-End Agent Systems through Persistent Real-Time Competition

Abstract

When users delegate work to an LLM agent, they care about both the quality of the result and how quickly it is delivered. Longer reasoning or additional agents may improve the result, but also introduce delays and coordination overhead. Therefore, an end-to-end evaluation should capture the consequences of these choices together while allowing systems to organize their own work. We introduce GeneralsArena, which uses real-time strategy competition as a controlled proxy for this outcome-oriented evaluation. Two systems explore a partially observed map and command forces to capture the opposing Capital while defending their own. Each system decides how to plan, act, and organize its agents, while the game continues during reasoning and communication. Further deliberation or delegation can help, but can also allow an opponent to seize an opportunity. Match outcomes jointly reflect decision quality, execution efficiency, and coordination. A persistent ladder compares complete systems, with subscription-service availability measured under the evaluation workload. Across 16 local runtimes, 96 runtime–harness–prompt combinations, and 24 subscription-service configurations, we find that systems sharing model weights can perform very differently, gains depend on the full system configuration, and service availability can change which service ranks first. GeneralsArena provides an evaluation framework for studying how system-level choices translate into end-to-end outcomes. Results, replays, and an introductory video are available at generalsarena.github.io.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.