acceptodds
Under review as a conference paper at ICLR 2027

DiLaT: Semantic-Guided Volumetric Diffusion for Through-Plane Coherent Brain T1 Denoising

Abstract

Brain T1 magnitude images are typically stored after reconstruction, where thermal noise follows a Rician law rather than a Gaussian one. Most learned denoisers still treat each axial slice as an independent 2D image. That is convenient, but it leaves a gap: when a volume is scrolled, independently restored slices can flicker even if per-slice SSIM looks fine. We treat a short stack of contiguous slices as the restoration unit and try to keep anatomy stable along the through-plane axis, not only sharp in-plane. DiLaT does this with a latent diffusion model: a frozen VAE provides the working space, frozen DINO features extracted from the noisy magnitude supply structure when pixels are unreliable, and a spatiotemporal Transformer exchanges information across the clip. On held-out AOMIC volumes the clip model reduces through-plane inconsistency relative to strong slice-wise CNN, Transformer, and diffusion restorers, while remaining competitive on SSIM, especially at higher noise. A zero-shot OASIS-1 test follows the same pattern. Ablating cross-slice attention or dropping the semantic condition at sampling both hurt volumetric coherence, which is the claim we care about.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.