SimStage: Harnessing Video Diffusion for Consistent Virtual Light Stage Synthesis
Abstract
One-light-at-a-time (OLAT) prediction provides explicit and reusable control for portrait relighting, but its promise depends on recovering an entire lighting basis with both Basis Fidelity and Basis Coherence: each response must preserve accurate HDR appearance, while all responses must remain structurally aligned so that details survive composition. We introduce SimStage, a Full-Basis OLAT Generator that directly predicts all lighting responses within a single generation window. Our key insight is Light-Axis Rebinding: rather than treating OLATs as independent predictions, we redesign the joint-generation paradigm of video diffusion so that its sequence axis represents illumination instead of time. This is non-trivial—native temporal position modeling fails to reliably learn even substantially shorter OLAT windows. We therefore introduce LiRoPE, which replaces temporal relative positioning with calibrated light-direction geometry and enables full-basis joint generation. Together with an HDR radiance encoding, SimStage produces an integration-ready OLAT basis from a single portrait. Experiments on FaceOLAT improve composited relighting PSNR by  dB over 3DPR, while preserving substantially sharper details and accurately controlling shading, cast shadows, and specular highlights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.