PriorDiffuDINO: Self-Trajectory Denoising for Diffusion-Based Detection Transformers
Abstract
We construct the diffusion process of a detection transformer as a denoising trajectory that starts from the noised encoder proposals and extends along the model's own predictions. We call this self-trajectory denoising and instantiate it on DINO as PriorDiffuDINO. The starting point of the trajectory does not depend on the ground truth and is shared by training and inference. Because the decoder is also trained on re-noised copies of its own predictions, which are the inputs of the later inference steps, the exposure bias between training and inference is reduced. We further adopt the matchability-aware loss and dense augmentation of DEIM, which also apply to diffusion-based detectors. With a ResNet-50 backbone and 12 training epochs, PriorDiffuDINO reaches 51.3 AP on the COCO 2017 validation set, 2.7 AP above DiffuDINO under the same schedule and close to DiffuDINO trained for 50 epochs (51.9 AP). With a ResNet-101 backbone it reaches 52.4 AP. On LVIS with ResNet-50, it substantially improves every metric over DINO and DiffuDINO, by 3.0 AP overall and 5.3 AP on the rare categories over DiffuDINO.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.