acceptodds
Under review as a conference paper at ICLR 2027

ProPAI: Proactive Physical AI for Perceiving, Reasoning, and Acting in the Open World

Abstract

Physical AI is commonly formulated around explicit tasks: a system receives an instruction, perceives its environment, and acts. In everyday settings, however, needs often emerge before anyone formulates a request. We study this setting as Proactive Physical AI, where a system continuously observes its environment, identifies needs under standing responsibilities, and initiates appropriate interventions. We introduce ProAgent, an end-to-end system that separates whether and when to intervene from how to respond. A Proactive Timing Model monitors the evolving scene, while an asynchronous Response Model generates communication or invokes available device capabilities, allowing observation to continue during response generation and execution. An automated data pipeline converts videos into causal supervision for learning intervention timing, and complementary inference optimizations reduce the cost of repeated decisions. Across streaming video evaluations, ProAgent achieves stronger intervention timing and time-aware response quality than the evaluated standalone multimodal models and streaming assistants under both explicit objectives and standing roles. Real-robot evaluations further demonstrate improved task completion across guidance, warning, and interaction scenarios. These results support a practical approach to physical AI that connects continuous observation with timely assistance and physical execution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.