acceptodds
Under review as a conference paper at ICLR 2027

DEYES: Denoise Everything You Ever Saw — Pretrained Visual Priors for Spatiotemporal Monte Carlo Denoising

Abstract

Physically based rendering produces much of the synthetic data that vision models train on, and Monte Carlo sampling sets its cost. A denoiser must reconstruct unbounded linear HDR radiance corrupted by heavy-tailed noise, a problem unlike photographic restoration. We ask whether visual representations learned from photographs transfer to Monte Carlo denoising. DEYES reads the features of a pretrained DINOv3 ConvNeXt encoder out as a multi-scale kernel pyramid that filters linear radiance with non-negative weights. Fine-tuned only on the open Noisebase corpus, it beats OIDN 2.5 and OptiX 9 on PSNR at every sample count from 1 to 512 on 39 scenes from an unseen renderer (Mitsuba 3), and beats OIDN, OptiX and NPPD on every temporal metric on Noisebase video. With its trunk frozen, DEYES beats every baseline on mean PSNR from one eighth of the training data and 9 A40-hours. Controls attribute the gain to the prior: the identical network trained from scratch peaks 1.04 dB lower and never reaches the best production denoiser on mean PSNR; a ConvNeXt pretrained by masked reconstruction on ImageNet-1k recovers most of DINOv3's gain, while DINOv3's own weights with their structure shuffled recover little. Measured end to end, DEYES reaches the best baseline's PSNR with 27–38% less render-plus-denoise time across quality targets from 32 to 512 samples per pixel. Two variants run at about 3 and 10 frames per second at 1920×1024 on one A40. Code, trained models and the Mitsuba evaluation set will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.