EgoSCOPE: Seeking What Is Missing via Query-Sufficient Evidence Routing for Long-Horizon Video Understanding
Abstract
Egocentric video understanding supports episodic memory, personal assistance, and embodied reasoning. In long-horizon recordings, answering a query requires combining observations scattered across time, viewpoints, and changing object states under a limited evidence budget. Existing approaches build long-video representations and structured memories or select query-relevant observations through ranking, diversity modeling, and adaptive search. However, individually relevant or diverse observations may still leave essential query relations unsupported, while repeated evidence consumes the remaining budget. Inspired by human problem solving, we propose Egocentric Spatiotemporal Closure over Progressive Evidence (EgoSCOPE), a fixed-budget router that progressively seeks evidence for unresolved query demands. EgoSCOPE represents each query's evidence requirements as a transient Query Demand Graph composed of query-specific relations from explicit semantic categories. A lightweight set encoder recurrently updates relation support and conflict, using marginal closure and necessity–sufficiency signals to prioritize unresolved demands and suppress redundant or conflicting evidence before a single frozen task-head call. Across Ego4D NLQ, EgoLifeQA, and VSI-Bench, EgoSCOPE consistently improves temporal localization, long-term memory, and spatial reasoning over representative task-specific systems and matched evidence-selection baselines, while increasing relation closure and reducing evidence duplication. These results demonstrate that explicitly tracking unresolved query demands enables more sufficient evidence composition under fixed budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.