acceptodds
Under review as a conference paper at ICLR 2027

MetaHuman-Anything: Vision-Language Face Similarity Grounding for Engine-Friendly Digital Human Sculpting

Abstract

Automated creation of editable 3D facial assets is increasingly important for digital human production in games and animation. However, image-based 3D face reconstruction methods typically output point clouds, meshes, or parametric face models rather than native characters that can be edited directly in animation engines. MetaHuman's mesh-conforming tools fit a native face template to reconstructed meshes, a step that can alter facial geometry and reduce identity fidelity. We present MetaHuman-Anything, an automated, vision–language-model-guided framework that directly edits native MetaHuman parameters to match a reference portrait. At each iteration, a frozen vision–language Reasoner identifies differences between the reference and the current render. A state-conditioned Grounder translates this analysis into grouped parameter targets and gated updates, while a learned renderer provides feedback for iterative refinement. We construct a corpus pairing rendered faces with editable MetaHuman states to train the framework. On an external set of 102 portraits, our method improves mean target-to-Unreal-Engine-render ArcFace identity similarity by 118.2% over the strongest mesh-conversion baseline under matched fixed-appearance conditions, while producing directly editable native MetaHuman characters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.