Realtime-Bench: Can Frontier Models Solve Real-time Interactive Problems?
Abstract
Real-time interaction requires agents to perceive evolving environments and act before opportunities disappear, yet existing evaluations for real-time streaming models focus on conversational quality and streaming perception, leaving their ability to act effectively under strict temporal constraints untested. We first formalize real-time capabilities for agents and introduce Realtime-Bench, a suite of dynamic game and web tasks in which delayed actions directly affect success. We evaluate six frontier LLMs spanning streaming models and standard asynchronous models under three clock protocols at low and high reasoning effort. Wall Clock advances environments with actual inference latency, while Pause Clock removes inference delays as an idealized reference. We further introduce Token Clock, which advances simulated time by the generation tokens consumed before action commitment, testing temporal awareness independently of physical serving speed. Best success rates reach only 32.0% under Wall Clock and 54.0% under Token Clock, compared with 84.0% under Pause Clock, revealing a substantial gap between task competence and timely action. Higher reasoning effort yields mixed effects under time pressure, reflecting a trade-off between decision quality and timely execution. Based on these results, we discuss future directions for improving real-time multimodal models, including adaptive reasoning with temporal awareness and continuous perception and action.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.