PBRSense: Single-Image PBR Estimation with VLM-Conditioned Image Editor
Abstract
Recovering the PBR (Physically based rendering) material properties of an object from a single image is a long-standing problem in computer vision and graphics. Traditional methods typically require multi-view inputs or controlled capture conditions and may take hours to solve the inverse rendering problem. Recent data-driven methods based on diffusion models can produce plausible material maps for in-distribution inputs but often struggle when the input falls outside the training distribution. VLM-conditioned image editing models provide a pretrained image-conditioned generation interface with language-controlled target specification. Whether such models can be adapted to predict PBR maps for general objects while addressing these limitations has not been systematically studied. In this paper, we propose PBRSense, which adapts a pretrained VLM-conditioned image editing model for pixel-aligned prediction of albedo, metallic, roughness, and surface normal maps from a single object image. Three attribute branches share the DiT backbone and use different text prompts to specify their prediction targets. We freeze the inherited VLM and VAE, selectively fine-tune DiT blocks, and use an endpoint velocity objective for single-step prediction. We compare trainable DiT ranges under a fixed training budget to inform the fine-tuning configuration. On the combined evaluation of two held-out benchmarks, the adapted model without communication already outperforms the evaluated baselines across the reported attributes; communication provides smaller additional gains. Qualitative results on real photographs illustrate its use beyond rendered inputs. Finally, using the predicted material maps as priors improves the material estimates of a frozen 3D generation model without additional training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.