ProtoDN: Appearance-Aware Denoising with Dynamic Visual Prototypes for Monocular 3D Object Detection
Abstract
Monocular 3D object detection is inherently challenging due to the depth ambiguity and incomplete observations from a single view. Recent denoising-based monocular 3D detectors focus on improving query learning and training stability. However, their denoising strategies remain largely centered on distribution modeling and geometric guidance, leaving instance-level visual information underexploited and limiting the modeling of challenging objects under occlusion, long-range observation, and truncation. To address this issue, we propose \method, an appearance-aware denoising framework with dynamic visual prototypes. We first introduce a dynamic visual prototype memory to aggregate class-specific visual priors. Building on this memory, we develop two core modules: (i) Geometry–Visual Joint Difficulty Modeling integrates geometric and visual cues to estimate instance difficulty and modulates denoising perturbations accordingly, enabling visually informed difficulty-aware denoising. (ii) Prototype-Guided Visual Enhancement leverages visual prototypes as instance-relevant appearance priors and explicitly injects them into denoising representations, compensating for insufficient visual information during reconstruction. Extensive experiments demonstrate superior performance of our method over recent approaches, with particularly notable improvements on challenging instances.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.