acceptodds
Under review as a conference paper at ICLR 2027

AvatarPilot: Learning a Camera Policy for Active 4D Human Avatar Reconstruction

Abstract

Capturing a moving person for 4D reconstruction typically relies on studio rigs with dozens of fixed cameras, which are costly and confine capture to the studio. An autonomous camera that moves around the person would be a far more flexible and cheaper alternative, but the quality of the reconstruction then depends on where the camera goes. We formulate this as active 4D human avatar reconstruction, a task in which a physically constrained camera decides at every step where to move next using only its onboard observations, to achieve a better 4D reconstruction from the images it captures. We solve this task by learning a camera policy with reinforcement learning. To train it, we turn hundreds of existing multi-view captures of real people into renderable 4D avatar assets made of 3D Gaussians, and build on them a pipeline from simulated onboard observations to camera actions, along with an evaluation protocol. We define the training reward directly on the avatars' surface Gaussians, which encourages the policy to see the body from new viewing directions. Experiments show that on held-out subjects, the learned policy outperforms every preset trajectory we compare against. When the person walks, the reconstructions from its captures exceed those from the strongest preset trajectory by 2.5 dB PSNR. When the person stays in place, one camera under our policy matches the reconstruction quality of six to eight fixed cameras.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.