GlassHead: Learning 3D Eyeglass-Wearing Heads from a Single Portrait via Pseudo-Multiview Supervision
Abstract
Single-image 3D reconstruction of eyeglass-wearing heads is challenging because important structures, particularly the temples and their attachments around the ears, are poorly observed from a frontal portrait. Learning to recover these regions requires diverse viewpoint supervision, yet large-scale multiview captures with varied eyeglasses are expensive and difficult to obtain. We address this data bottleneck by using generative video as a scalable source of pseudo-multiview supervision. We introduce a scalable pipeline that leverages a video generation model to transform frontal portraits into view-diverse pseudo-multiview observations, recover their camera parameters, and construct the Eyeglass Portrait Pseudo-Multiview Dataset (EPPM-20K). Building on this supervision, we propose GlassHead, a feed-forward reconstructor that predicts a complete 3D Gaussian head from a single portrait through hierarchical coarse-to-fine decoding. To accommodate inconsistencies in generated views, we combine reliable frontal reconstruction with perceptual and camera-conditioned adversarial supervision on pseudo-multiview observations. Experiments demonstrate consistent improvements over recent single-image head reconstruction methods, with particularly clear gains in lateral eyeglass structures under novel viewpoints. At inference time, GlassHead requires only a single frontal portrait and reconstructs the 3D head in one forward pass within one second.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.