GOLD: Generalized Proposal Denoising for Domain-Generalized Object Detection
Abstract
Domain-generalized object detection (DGOD) aims to train a detector solely on labeled source domain data that generalizes directly to unseen target domains without accessing target data or performing test-time adaptation. Despite encouraging progress, existing approaches largely optimize encoded features or final predictions within deterministic detection frameworks, and thus lacking an explicit mechanism to iteratively recover object locations from unreliable spatial hypotheses. To this end, we propose **G**eneralized Pr**O**posa**L** **D**enoising **(GOLD)**, a diffusion-based DGOD framework that reformulates cross-domain localization as a semantics-preserving denoising process that progressively refines noisy proposals under unreliable visual evidence. First, a Spatial Calibration Adapter (SCAdapter) equips frozen DINOv3 features with localization-sensitive spatial structure for iterative proposal refinement. Second, Spatially Heterogeneous Evidence Perturbing (SHEP) decomposes visual evidence into stochastic foreground regions, localization-critical contexts, and distant backgrounds, as well as applies moderate, strong, and weak residual perturbations, respectively, according to their roles in proposal denoising. Third, Prototype-Anchored Denoising (PAD) constructs multi-scale category prototypes and anchors matched proposal representations at every refinement stage, preserving category semantics throughout proposal denoising. Extensive experiments on multiple cross-domain benchmarks demonstrate that GOLD consistently outperforms previous DGOD methods. *Our code is released in supplementary materials.*
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.