acceptodds
Under review as a conference paper at ICLR 2027

Streaming 4D Hand Interaction Forecasting with an Autoregressive Video Model

Abstract

We present a streaming 4D interaction forecaster that turns an autoregressive video model into a structured predictor of future hand motion and contact. Given recent egocentric visual and state observations, the model predicts future interaction as a joint state of bimanual articulated hands, egocentric camera motion, and joint-wise hand-object proximity. This structured representation captures both motion and interaction cues in explicit geometry, providing a forecasting space that is more directly usable than generated video alone. We realize this representation with a 4D state decoder that extracts structured interaction trajectories from intermediate denoising features of an autoregressive video model. The decoder uses current hand and camera states to ground its predictions in the present scene, while incoming observations continually update the forecast. Trained on egocentric clips spanning lab-collected 3D data and diverse in-the-wild videos, our method improves articulated hand forecasting over prior baselines under matched observation and state access, while additionally predicting future hand-object proximity cues. We further demonstrate contact-guided interaction in physics simulation using the predicted hand motion and proximity cues as control references.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.