Bootstrap Your Own Noise: Denoising and Latent Prediction Cooperate Only When Noise Is Informative and Alignment Is Gated
Abstract
Reconstruction and latent prediction teach a vision encoder different things. Reconstruction from masked or noised inputs yields features that attend locally. Predicting the embedding of a clean view yields features that attend globally, so combining the two is a natural goal. We show that the combination is not additive. We pre-train a ViT-B for 400 epochs on ImageNet-1K. Adding Gaussian latent noise to masked image modeling leaves fine-tuning accuracy unchanged (82.86% against 82.89%), and adding a bootstrapped latent-prediction loss on top lowers it to 82.01% and costs 4.6 mIoU on ADE20K. We trace both failures to a single cause and remove it with two small changes. Image-adaptive noise sets the noise level of each image from its token saliency, and a reliability gate scales the latent-prediction loss by the current reconstruction error. The gate alone returns accuracy to 82.84%, adaptive noise alone reaches 83.16%, and the two together reach 83.56%. In other words, the latent-prediction loss helps only after the noise carries information about the image and the target it aligns to can be trusted. Under the same budget, the resulting model, Bootstrap Your Own Noise (BYON), improves over masked and denoising pre-training by 0.7 points on ImageNet-1K after fine-tuning and by 2.2 under a linear probe, by 1.4 mIoU on ADE20K, and by 1.7 box AP on COCO. It gains 0.8 to 3.5 points on six fine-grained benchmarks, and its margin over MAE does not shrink at 800 epochs or with ViT-L. Our evidence is empirical and covers models up to ViT-L trained on ImageNet-1K; code is provided in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.