acceptodds
Under review as a conference paper at ICLR 2027

Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

Abstract

Proactive multimodal assistants must decide guidance is needed and to say, while tracking progress through extended procedures and helping users recover when they depart from the expected plan. Progress is limited by benchmarks that score isolated clips, cover narrow domains, or omit realistic deviations. We introduce , a smart-glasses dataset of procedures across four domains in which users commit step-level deviations and recover from them, paired with reference guidance. We combine it with five re-annotated egocentric and instructional datasets to form , a cross-domain benchmark that scores intervention timing, procedure completion, and response quality. Zero-shot, no frontier or open-weight model completes more than 26% of a procedure's interventions before its first error. We propose (Plan, Watch, Recover), which keeps an explicit plan state in the model's context and revises it after every intervention, including recovery steps after a deviation. Across three open-weight backbones, training on without a plan barely changes performance, while adding the dynamic plan raises average procedure completion from 20% to 27% and improves turn-level decisions and response quality for every backbone. The gains hold on held-out datasets and on deviations never seen in training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.