Preemptive Instruction Following With Re-Sumption
Abstract
Recent advances in embodied AI have enabled agents to follow natural languageinstructions for navigation. However, real-world interactions often involve dy-namic user behaviors such as issuing instructions before the agent is ready orinterrupting and later resuming tasks, which existing models struggle to handle.To bridge this gap, we introduce Preemptive Instruction Following with Resumption (PIFR), a novel paradigm that requires agents to interpret instructions givenat arbitrary times and seamlessly resume previously interrupted goals. To sup-port research in this direction, we construct a new multi-task dataset coveringobject-goal navigation and vision-language navigation, augmented with realisticpreemptive and resumption scenarios. We conduct comprehensive evaluations of awide range of monolithic and modular methods, analyzing their capabilities in in-struction buffering, maintaining goal context under interruption, and generalizingacross task types and environments. Our benchmark reveals significant limita-tions in current approaches and provides a foundation for developing more robust,human-aligned interactive navigation systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.