acceptodds
Under review as a conference paper at ICLR 2027

Generative Lighting Estimation for Multi-View Captures

Abstract

An environment map is the standard representation of the distant illumination arriving at a point in a scene, and many inverse rendering and relighting pipelines on real captures depend on one. Generative methods predict a map from a single image with a learned prior. Everything outside that image's field of view is then inferred from indirect cues rather than observed, so the estimate does not faithfully reproduce what is in the scene. Per-scene inverse rendering instead optimises against a whole multi-view capture, solving for illumination jointly with geometry and materials. It is constrained only where the capture has observations, and the light sources usually lie outside them. We recover a high dynamic range environment map from a sparse set of posed low dynamic range frames, so that the learned prior supplies only what the capture does not. The map is produced as a set of views that together cover all directions, generated jointly by a multi-view diffusion model so they agree with each other and with the input frames, and assembled into a panorama. The single-image setting is the smallest case of this. On it we improve on the state of the art on the Poly Haven and Laval Indoor benchmarks. Additional frames improve the estimate further, with diminishing returns. On captures with a moving camera, the frames need not be taken where the map is centred; the benefit falls off slowly with distance. Each frame enters with its pose, so unlike stitching and outpainting the model is not bound to a single viewpoint. A probe over scenes with near-field landmarks confirms this.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.