acceptodds
Under review as a conference paper at ICLR 2027

LeDenoiser: Exploring Latent World Models for Monte Carlo Rendering Denoising

Abstract

Monte Carlo rendering converges as samples accumulate, yet deep learning-based denoising rarely models that process. The field has largely settled on image-to-image regression: spatial 2D filtering over rigid multi-channel inputs, supervised only on output radiance. We ask whether light transport convergence can instead be formulated as an evolving physical state transition. Inspired by latent world models, which learn action-conditioned dynamics entirely within a regularized representation space, we propose *LeDenoiser*, a latent world model for Monte Carlo denoising. A noisy 1-SPP observation forms the initial state, encoded by a pre-trained vision transformer; auxiliary Arbitrary Output Variables (AOVs) act as discrete token-level action operators; and the predictor learns an action-conditioned transition that advances this state toward convergence. On SPP-World, our 22-scene multi-SPP path-tracing benchmark, unconditional latent dynamics lift severe 1-SPP noise (17.75 dB) to 32.61 dB, demonstrating that a latent world model can capture scene-level light transport without treating auxiliary geometry as a rigid dependency. Because AOVs form a discrete action space, that space can also be planned over. To our knowledge, this provides the first planning study inside a rendering world model: Model Predictive Control via the Cross-Entropy Method discovers sparse action subsets ( of 8 buffers) that outperform exhaustive all-buffer conditioning. The latent dynamics are controllable, and geometric redundancy varies non-trivially across scenes. These findings establish latent world models as a viable alternative to image regression in Monte Carlo rendering, bridging graphics pipelines with physical foundation models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.