STANCE: Scene-Consistent Real-Time Motion Estimation from Sparse Wearables via Posterior-Corrected Residual Diffusion
Abstract
Wearable tracking began as a way to drive game avatars and has since become an instrument for embodied intelligence, where human demonstrations captured from a few body-worn sensors train agents inside simulated environments. Configurations vary widely: a headset and two controllers give three 6-DoF signals, and optional hip and foot trackers extend this to six. Methods that accept arbitrary sensor sets fix each tracker's influence in advance, zero-padding or masking absent channels with hand-crafted weights that encode measurement presence rather than reliability. A second gap opens only inside the scene: signals captured in a physical room are replayed in a virtual one, and the mismatch surfaces as limbs passing through the very geometry the motion is meant to interact with. We present STANCE, which predicts and then corrects. A configuration-agnostic predictor forecasts a temporally coherent prior from a fixed set of core sites, so continuity never depends on which trackers are present, while a single-frame residual denoiser rewrites only what the trackers contradict, keeping accuracy and latency from trading against each other. How far a partial inverse-kinematics solve travels from the prior, and how closely it meets the trackers, give each joint a continuous trust gating how much the denoiser may rewrite; a configuration change then alters feature values and nothing else. Inside the denoiser, a training-free posterior-sampling procedure corrects the predicted residual against scene geometry at every diffusion step, removing penetration without retraining or a learned collision model. STANCE runs in real time, reaches state-of-the-art accuracy across most tracker counts, and holds continuity through dynamic reconfiguration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.