acceptodds
Under review as a conference paper at ICLR 2027

RTWorld: Benchmarking Real-Time Game Playing Agent for Open-Ended Tasks in Digital Environment

Abstract

Generalist visual agents must perceive, plan, and act while the digital world continues to change, yet most evaluations implicitly let the world wait for model inference. This convention obscures whether an agent can turn a correct decision into a timely action. We introduce RTWorld, a benchmark for real-time game-playing agents built from 31 open-ended browser games with pixel-only observations, mixed keyboard and pointer control, and verifiable native outcomes. A fail-closed runtime preserves the live world clock, timestamps the full observation–inference–action path, rejects invalid or stale plans, binds each episode to one model replica, and records replayable hash-verified evidence. We evaluate 13 vision-language systems using common seeds and complement the main benchmark with controlled-delay, pause/live, delay-shape, and temporal-context interventions. Our results show that static model quality, response availability, and closed-loop action freshness are distinct factors: model rankings can change when inference no longer pauses the environment, and faster responses do not by themselves guarantee useful control. RTWorld provides a reproducible foundation for studying agents that must reason and act before the opportunity has passed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.