LEXIS: Latent Proximal Interaction Signatures for 3D HOI from an Image
Abstract
Reconstructing 3D Human–Object Interaction from an RGB image is essential for perceptive systems and it requires capturing the subtle physical coupling between the body and objects. Current methods rely on sparse, binary contact cues. These miss the continuous proximity and dense spatial relationships inherent in interactions. We address this with InterFields, which encode dense, continuous proximity and the direction to the interacting geometry across the entire body and object surfaces. Inferring such fields from a single image is ill-posed. Our intuition is that interaction patterns are structured by the action and object geometry. We capture this structure in LEXIS, a novel discrete manifold of interaction signatures learned via a quantized autoencoder. We then develop LEXIS-Flow, a flow-matching framework that uses LEXIS manifold to estimate human and object meshes alongside their InterFields. These InterFields guide sampling that improves physical plausibility without post-hoc optimization. On Open3DHOI and BEHAVE, LEXIS-Flow outperforms SotA baselines in reconstruction, contact and physical plausibility. It generalizes better and the reconstructions are perceived as more realistic, moving us closer to holistic 3D scene understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.