Label Noise Forgotten in Loss Persists in Complexity
Abstract
Training loss can decline as a network memorizes incorrect labels. What remains observable about label noise after its initial loss response subsides? We identify an asymmetry in training trajectories: after controlled label-noise injections, loss returns toward its pre-injection course, while neural persistence, a statistic of normalized network weights, remains elevated through the recorded horizon. Matched clean additions do not produce the same sustained elevation. Separately, clean and noisy per-sample losses overlap strongly on real-world image data with injected label noise, despite clear separation on a simple synthetic dataset. We study a geometric mechanism for the complexity cost of fitting noise: maintaining a fixed positive output separation across nearby conflicting inputs requires increasing functional complexity. Explicit calibration and trajectory conditions connect this barrier to measured persistence and label recovery. We use the persistent signal in our Concept Recovery Test (CRT), combining classwise label-fit weights with a global complexity-level gate to aggregate training predictions. We prove finite-time recovery when the gate retains enough clean predictions, accounting for early errors, reference estimation, and the memorized tail. Across five benchmarks and four comparison methods, gated CRT improves rectification and detection over both the underlying method and loss-only CRT, and remains effective as oracle-tuned single-epoch loss thresholds deteriorate. These results show how a sustained structural response can complement loss-based evidence for label recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.