Beyond Latency: Benchmarking Real-Time Strategic Control for LLM Agents in Evolving Environments
Abstract
Large language models are increasingly used in Agent tasks that require continuous observation, reasoning, and action. In real-time environments, however, the world continues to evolve during model inference. Each decision therefore spans both wall-clock time and environment time, affecting the decision frequency achievable within a fixed horizon. We develop an OpenRA-based benchmark for real-time strategic control, comprising controlled Fixed-Force diagnostics and complete Full-Skirmish games. The benchmark jointly records wall-clock response time, observation-to-action gap, decision frequency, and control outcomes. Experiments show that added waiting widens the gap while reducing decision frequency. The same wall-clock waiting time corresponds to substantially different gaps under different world speeds; the mean gap at the fastest speed is roughly three times that at the slowest. Full-Skirmish experiments further reveal clear differences in computational cost and long-horizon control performance across observation representations and control architectures. The benchmark also supports evaluation across model backbones, Bot behaviors, and opponent difficulties in continuously running RTS environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.