ReUP: Retrieval-Augmented Understanding and Prediction of Human Behaviour from Sparse Egocentric Signals
Abstract
An assistant such as smart glasses needs a holistic picture of its wearer. It must know what the wearer does, where the wearer is, how the body moves and what comes next. Existing systems run large models on continuous egocentric video or train one model per question. We show that two cheap signals suffice: the head pose the glasses already track and one image. We present ReUP, a Retrieval-augmented model for Understanding and Predicting human behaviour. Its memory holds fully instrumented walkthroughs, where every moment carries all the answers. Multi-modal contrastive learning gives its head-motion key rich semantics from the body and narration of each moment. Behaviour-guided contrastive learning draws together moments that share a place, an activity or a future. At query time, head motion and image each retrieve similar moments. For each task, a learned scorer picks the useful ones from both, and a light decoder turns them into an answer. The model blends what only each cue offers: the place from the image, the activity and body from the motion. With fewer than 2M parameters, one model recognises the current and next action, estimates and forecasts the body pose and localises the wearer. In unseen buildings and on an unseen dataset, it is on average 3.5% better than the strongest baseline of each metric. Retrieving recorded human behaviour gives a lightweight model a prior for holistic and effective understanding and prediction, from minimal input and in unseen scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.