RTVA: A Framework for Real-Time Visual Agents in Evolving Environments
Abstract
Alongside task-oriented visual agents that operate through turn-based action loops, recent proactive visual agents aim to support more natural interaction by continuously perceiving evolving environments and responding in real time. Enabling such agents requires more than low latency: perception, response generation, and LLM reasoning proceed concurrently but at different timescales, making it difficult to keep ongoing interaction aligned with a changing environment. Different tasks also demand distinct patterns of intervention and adaptation. We characterize streaming visual interaction through three temporal boundaries: before response initiation, during response delivery, and when delayed reasoning returns. We introduce RTVA, a framework that decomposes real-time visual interaction along these boundaries into composable components within a dual-system architecture. These components separate content generation from whether, when, and how generated content may influence interaction. To adapt RTVA to task-specific interaction requirements, AutoHarness searches component compositions across these boundaries, then optimizes their behavioral policies. RTVA achieves the strongest overall cross-domain performance in both temporal alignment and interaction quality for real-time visual interaction. We further introduce the Visual In-Flight (VIF) Benchmark, a complementary capability evaluation of whether ongoing responses remain grounded, current, and coherent as visual evidence changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.