acceptodds
Under review as a conference paper at ICLR 2027

EgoDual: Multi-Level Task Planning from Egocentric Video

Abstract

A system that guides a procedural task from video, whether a proactive assistant coaching a person or a planner instructing a robot, must decide not only what comes next, but when to intervene. The right level of guidance depends on who carries it out: a person can act on the next subtask, while a robot policy needs the next concrete action. These levels follow different timelines. Within a subtask, the next action is imminent while no subtask boundary is near. At a boundary, the opposite holds. A granularity-agnostic policy cannot be correct at both. We introduce granularity-conditioned procedural planning, in which the requested level (subtask or action) is an explicit input that determines both what the agent predicts and when it intervenes. We build hierarchical supervision over subtasks and actions, and introduce EgoDual, a dual benchmark of 7,017 queries from 93 egocentric procedural videos in which coarse and fine requests share the same visual input. Across four VLMs, zero-shot models are largely insensitive to granularity. Our dual supervision learns distinct coarse and fine intervention policies, raising the granularity separation score from 0.45 to 0.68, and improves action-level prediction. Our analysis further shows that subtask-level anticipation is bottlenecked by predicting the next procedural state, not by generating text at the right level of abstraction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.