LOOM: Low-latency Object Onset Modeling for Egocentric Audio-Visual Anticipation
Abstract
For wearable assistive systems designed for people with visual impairments, timely alerting users to surrounding hazards, critical events, and key objects is paramount to ensuring safe navigation and independent mobility. We formulate egocentric audio-visual near-body intrusion anticipation: predicting whether an object will newly enter a region extending 0.40m beyond the body within three seconds, using historical monocular RGB and simulated two-channel audio. We introduce LOOM, a lightweight framework combining visual motion and expansion statistics, a learned acoustic sequence model, development-selected probability fusion, and stateful warning generation. We construct ReactionGap, comprising 150 simulated episodes across ten motion families, with fixed wearer-centered viewpoints, synchronized audio-visual observations, and geometric trajectories. Its entering and non-entering scenes provide 3,900 labeled queries for training and template-disjoint evaluation of future-entry prediction and sequential warnings. Across three seeds, LOOM achieves pooled average precision of 0.5438, compared with 0.5234 for vision and 0.4776 for audio. Median computation takes 8.8ms, compared with 54.1–72.4ms for the evaluated temporal and fusion adaptations. Warning replay yields one-second recall of 0.4178 at 0.7289 non-useful alarms per episode; vision retains stronger time-matched discrimination. These results highlight a critical discrepancy between pure model prediction and actionable warning performance, establishing ReactionGap as a standardized benchmark to advance robust, real-time wearable assistance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.