acceptodds
Under review as a conference paper at ICLR 2027

CUAWorld-RT: Can Vision-Language Agents Think and Act in Real Time?

Abstract

AI agents have made substantial progress in domains where they have a non-trivial amount of time to think between actions, such as mathematics, software engineering, and performing tasks on desktop applications. However, it is unclear how such models fare in settings that demand real-time interaction, such as playing a video game or operating a physical device. In these settings, the environment state may evolve as the agent is reasoning, or even as the underlying model is performing a forward pass. To gain insight into the real-time interaction capabilities of vision-language agents, we construct CUAWorld-Realtime (RT). Our key idea is to construct visual minigames that demand a portfolio of capabilities required for interactive real-time settings. We build an agentic pipeline for automatically generating such environments based on a taxonomy of six capabilities required for real-world and real-time interaction. comprises 100 minigame families and 1,000 configurations that vary across task difficulty and the action interface. A key feature of CUAWorld-RT is that it allows for controlled comparisons across different capabilities, such as real-time interaction, temporal reasoning, and exploration of the action interface. Using this, we evaluate frontier and open-weight models. Even GPT-6 Astra solves only 12.5% of real-time tasks at the highest difficulty, while being extremely slow. We also find that models' strong coding capabilities can partially offset their weak real-time capabilities, although their performance remains significantly low. We will publicly release our environment generation framework to spur further research on AI agents in real-time interactive settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.