HybridAvatar: Photoreal Animatable Humans as Mesh-Native Assets from Multi-view Cameras
Abstract
Mesh-based avatars have enjoyed decades of graphics engine support: physics, shadows, and relighting all assume mesh geometry. Recent 3D Gaussian Splatting (3DGS) avatars deliver higher rendering quality and can be reconstructed from multi-view captures with as few as 10–16 cameras, but their splat representation is incompatible with conventional graphics engines. We present a hybrid pipeline that bridges these two regimes. We train a pose-conditioned avatar with mesh-anchored Gaussian Splatting from multi-view video, then, for every training pose, extract three asset maps (texture, normal, displacement) with a UV-space rasterizer applied to the trained Gaussians. At inference time, geometry for a novel pose is produced by the trained network as a posed mesh, while appearance is produced by a pose-space mixture-of-Gaussians blend over the per-pose asset maps. The output is a textured mesh asset that can be loaded directly into any engine, supports physics and dynamic shadows, and responds to scene lighting, while approaching the rendering quality of 3DGS avatars. The same model can also be rendered directly as Gaussians at quality comparable to 3DGS avatars. In a hybrid mode it is rendered with Gaussians while its mesh, which shares their geometry, serves as a physics and shadow proxy. We evaluate the pipeline on multi-view human captures and show that it produces engine-ready assets whose photometric quality exceeds that of meshes extracted from 3DGS avatars and approaches that of direct renderings of animatable 3DGS avatars. Our project page is available at https://anonymous.4open.science/w/HybridAvatarICLR27/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.