SLoHa: A 4D Human Locomotion Dataset of Scene-constrained Large-object Handling via a Democratic Data Acquisition Solution
Abstract
Carrying large objects through cluttered environments is an important capability for robots in daily assistance, logistics, and workplaces, requiring coordinated whole-body motion, object control, and obstacle avoidance. Yet learning this capability remains underexplored, partly due to the scarcity of suitable training data. Human demonstrations provide a valuable reference for robots to learn such coordinated behaviors. However, existing datasets primarily capture small-object interactions in unconstrained environments and often rely on expensive motion-capture systems that are difficult to scale, limiting the availability of scene-aware large-object demonstrations and hindering data-driven learning of this capability. To bridge this gap, we present SLoHa, a dataset for scene-constrained large-object handling, together with a democratic data acquisition solution that makes such demonstrations more accessible to collect. Using four exocentric cameras, modern reconstruction priors, and sparse human-assisted alignment, our pipeline recovers human motion, object geometry and 6D pose trajectories, and scene constraints in a shared coordinate system without specialized MoCap suits or object markers. SLoHa contains 21.4 hours of demonstrations involving 150 large objects and 435 obstacle layouts. We present a scene-conditioned large object handling benchmark. Experiments show that our dataset improves constraint-compliant motion generation, validating its effectiveness for studying scene-aware interaction and its potential use in embodied applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.