acceptodds
Under review as a conference paper at ICLR 2027

SituHOI: Scalable Capture of Scene-Situated Full-Body Human-Object Interaction

Abstract

Modeling human-object interaction (HOI) in real environments requires jointly representing whole-body motion, manipulated objects, and scene geometry, yet existing datasets trade off scalable capture against structured, scene-aligned supervision. We introduce , a large-scale RGB-D dataset containing 18,429 annotated clips across 16 indoor rooms and approximately 86 hours of quality-screened activity. It pairs recordings from a single moving RGB-D camera with static scene scans, object geometry, language descriptions, and estimated human and object trajectories. Using the scans as persistent references, our pipeline combines camera localization and depth-based metric alignment with human mesh recovery, object tracking, and temporal refinement to register all geometric annotations in a shared floor-aligned frame. These data support , a goal-conditioned joint motion prediction benchmark with variable observation and prediction durations. Given observed histories, scene and object geometry, and a terminal object pose, models predict future human and object motion. We provide as a reference predictor and evaluate world-space accuracy and human-object coordination. On a MoCap subset, our pipeline reduces human G-MPJPE from 195.4 to 68.3 mm and object translation error from 43.0 to 28.7 mm. SharedHOI leads four adapted baselines on all three aggregate FlexHOI-Bench metrics. Data will be released upon acceptance. An anonymized code repository is available at https://anonymous.4open.science/r/iclr2027-anonymous-code-1DFB.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.