acceptodds
Under review as a conference paper at ICLR 2027

LightShift: Controllable Indoor Relighting via Video Diffusion Priors

Abstract

Controllable image-based indoor relighting aims to render a source image under user-specified illumination conditions while preserving intrinsic scene properties and structural consistency. Previous approaches typically leverage pretrained im- age diffusion models with auxiliary conditioning branches, such as ControlNet, to incorporate estimated geometry (e.g., depth and normals), intrinsic properties (e.g., albedo), and lighting instructions. However, errors in these estimates can propagate to the relit outputs, degrading visual quality and scene consistency. We present LightShift, a controllable indoor relighting framework that leverages pre- trained video diffusion priors without requiring geometric or intrinsic estimates at inference time. Our key idea is in-sequence visual conditioning: we repre- sent the source image, a scribble-based lighting instruction, and the relit output as a unified three-frame sequence. Rather than injecting the lighting instruction through an auxiliary conditioning branch, this formulation uses temporal interac- tions to integrate scene context and spatial lighting control when generating the relit frame. To support training, we introduce a two-stage data synthesis pipeline that constructs relighting pairs from in-the-wild indoor photographs. The first stage generates coarse relighting pairs through intrinsic-based re-rendering, while the second uses the first-stage model to synthesize relit source images paired with real photographic targets. Quantitative and qualitative evaluations, together with a user study, demonstrate the effectiveness of LightShift in controllable indoor relighting, with improved lighting control and scene consistency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.