AHA! Animating Human Avatars and Dynamic Objects in 3D Scenes with Gaussian Splatting
Abstract
We present AHA!, a framework that animates human avatars navigating and interacting with dynamic objects in 3D scenes, with 3D Gaussian Splatting (3DGS) as the only 3D representation for rendering. Existing animation pipelines rely on meshes or point clouds for scenes and humans. These provide the surfaces, distance fields, and walkable regions that motion synthesis and contact modeling need, but they do not provide photorealistic renderings and are hard to recover from casually captured video. 3DGS offers photorealism and reconstruction from a single mobile-phone video, yet provides none of these geometric quantities. We close this gap with four components that operate directly on opacity-filtered Gaussians: occupancy maps for path planning, an egocentric walkability grid that lets a locomotion policy trained only in synthetic mesh rooms run zero-shot in 3DGS scenes, a differentiable SDF proxy for synthesizing scene interactions, and a differentiable contact refinement for human-scene, human-object, and object-scene contact. Because maps and contact targets are functions of the current object poses, the same stack composes long animations in which an avatar walks through a scene, sits down, lifts objects, carries them, and puts them down, re-planning as its own actions change the layout. On SuperSplat and DL3DV scenes, human and VLM preference studies favor our renderings over the strongest mesh pipelines, and our motion quality is within 5% of a clean-mesh upper bound. Full-resolution videos are available at https://anonymous.4open.science/w/AHAICLR27/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.