acceptodds
Under review as a conference paper at ICLR 2027

EgoSkill: Rethinking What to Store and How to Reason in Long-Form Egocentric Video Question Answering

Abstract

Long-form egocentric video question answering (VQA) requires reasoning over evidence that may be sparse, distributed across distant moments, and heterogeneous across question types. We introduce EgoSkill, a fully training-free framework that factorizes this problem into two complementary inference-time components: what to store and how to execute. For what to store, we propose Dynamic Hierarchical Memory (DH-Mem), a query-agnostic representation of long-form egocentric video. DH-Mem represents local events as temporally grounded substeps and semantically composes related substeps, even when they are temporally non-contiguous, into higher-level procedural steps and a video-level goal. For how to execute, we introduce Evidence Construction Skills (EC-Skills), reusable query-aware procedures that adapt evidence retrieval, visual processing, tool use, and verification to the requirements of each question type. Together, these components realize a store-execute factorization, with DH-Mem providing query-agnostic video context and EC-Skills constructing query-specific evidence for frozen MLLMs, without task-specific parameter updates. Across HD-EPIC, EgoSchema, and QAEgo4D, EgoSkill consistently improves the corresponding frozen MLLM baselines, demonstrating the effectiveness of factorizing long-form egocentric VQA into query-agnostic representation and question-dependent evidence construction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.