EgoReView: Memory-Guided Active Visual Evidence Acquisition for Long-Horizon Egocentric Question Answering
Abstract
A personalized egocentric assistant has to recover fine details from a long video history and to keep the observations that will matter for later questions. Sparse frame overviews and caption memories tend to omit exactly the evidence a person asks about. We present , an agent harness that couples active evidence acquisition with memory self-improvement. Starting from up to 512 frames and an initial answer, the agent retrieves from memory, revisits video intervals, and inspects cropped regions as unresolved gaps demand. Only decisive supplementary evidence licenses a revision of the answer, and observations that survive verification against their source media enter personal memory for the following days. Neither the model weights nor the acquisition policy change. On EgoMemReason, reaches 54.2% accuracy, the best result among the methods we compare against, versus 52.2% for its direct baseline. On a seven-day, single-person EgoLifeQA stream, retaining verified observations raises the fixed agent from 50.2% to 51.8%, while the direct baseline reaches 48.4%. Refining memory through ordinary use is thus a practical route to personal adaptation, in which every interaction becomes an opportunity to improve the evidence available to the next question.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.