PixelBody: Multimodal Priors for Monocular Human Mesh Recovery
Abstract
Feed-forward human mesh recovery predicts 3D bodies from images and video in one shot but does not fit them to the input pixels, so the results can disagree with image cues such as landmarks, silhouettes, and surface normals. PixelBody is an offline method that starts from the body trajectory of a feed-forward method and iteratively fits it to such cues. During the fitting process, it utilizes a pose autoencoder pretrained on motion capture data. However, instead of applying it as a regularization, PixelBody adapts the decoder to each person: its weights are first fine-tuned to reproduce the initial trajectory from the encoded pose latents, and then the decoder and latents are fitted together to the video cues. Our ablations show that fine-tuning the decoder prevents the fit from being limited to the motion-capture prior, while still benefiting from it, and achieves lower 3D error than both direct pose optimization and a frozen decoder. Fitting runs in two stages: registration to skeleton and surface landmarks, then refinement that adds silhouettes, body normals, and relative depth. The landmarks, normals, and depth come from predictors we fine-tuned from Sapiens2 and Wan 2.1. PixelBody has the lowest camera-space error on EMDB1 among published methods that recover a body mesh, and, with ground-truth cameras, the lowest world-space error on EMDB2 among the compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.