Frozen Foundation Models and Noisy Labels: Accuracy Headroom and Calibration Gaps
Abstract
Foundation-model encoders change the starting point for learning from noisy downstream labels: a visual representation is available before the corrupted supervision is observed. We ask how much accuracy headroom remains for noise-specific training when the encoder is immutable, and whether accurate readouts also yield reliable probabilities. Linear probes serve as matched diagnostics over the same frozen representation, not as substitutes for representation learning. Across the evaluated benchmarks, cross-entropy probing often approaches a matched clean-label readout and specialized noisy-label learning rules under moderate noise, while severe or structured corruption leaves meaningful accuracy headroom. Absolute accuracy also varies across encoders, so this finding is conditional on the tested representations. Yet similar top-1 accuracy can coexist with markedly different calibration errors. A population analysis separates changes in Bayes decisions from changes in posterior probabilities. We then evaluate Anchor-based Efficient Calibration (AEC), a post-hoc temperature-matching intervention that uses training logits rather than clean calibration labels. AEC often improves ECE and representative proper scores without changing predictions, but its fixed target can also worsen calibration. These results position decision quality and probability reliability as distinct objectives in noisy-label learning with frozen foundation models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.