EgoEasy Annotator: Evidence-Governed Task-Navigation Timelines from In-the-Wild Egocentric Video
Abstract
Human egocentric video could scale robot learning, but an untrimmed recording is not a ready-made training example. Recovering tasks, operations, and movement between work areas requires reading fleeting or obscured clues without inventing details or missing events. We introduce EgoEasy Annotator (EEA), which checks proposed descriptions against the video and separately searches for missed events. When interpretations would change the annotation, it asks what the video must show to distinguish them and opens suitable new views. A video model describes these views without seeing the descriptions being checked. A controller retains a specific description when the evidence distinguishes it from the alternatives, or shared details when only those are supported. Unclear details that still affect the annotation receive local human review; unsupported guesses are omitted. Thus, uncertain object or work-area names need not erase supported actions or navigation. Each newly found event undergoes the same checks before entering one source-timestamped task–navigation timeline. Averaging HD-EPIC and EgoEasy-Bench scores, EEA improves episode segmentation F1 by 12.34 percentage points over the best baseline, while recovering more events with fewer unsupported claims. In a 36-video study, it saves 32.14% of human review and correction time over the development-selected baseline while meeting the preset requirements for average final quality across both benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.