RAD: Reconstruction-Anchored Diffusion Pretraining for Corruption-Robust Point Cloud Recognition
Abstract
Corruption-robust point-cloud recognition has become an essential task for the realm of 3D vision. Current globally conditioned diffusion pretraining typically optimizes the encoder for whole-shape denoising through a pooled global condition. Although effective in transferring features for clean recognition, this approach leaves masked-patch geometry without an explicit prediction objective. In this work, we propose a framework, named as RAD, for corruption-robust point-cloud recognition, via complementing whole-shape generative pretraining with an explicit geometric prediction task. Specifically, RAD consists of two key modules, the Local Geometry Reconstructor and the Conditional Flow Denoiser. The former module, the Local Geometry Reconstructor, is designed to reconstruct masked patches from visible encoder tokens under Chamfer distance, thus supplying the missing geometric prediction objective. The Conditional Flow Denoiser retains whole-shape denoising as an auxiliary pretraining objective, thereby keeping the generative branch. Through the integration of these two modules, RAD jointly pretrains a single encoder that is transferred alone, effectively leaving the downstream classifier and fine-tuning protocol unchanged. Extensive experiments across diverse corruption benchmarks demonstrate that adding geometric reconstruction reduces mean Corruption Error relative to the diffusion baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.