What Is Averaging Worth in Self-Training? Overlapping Correlation Bounds across Steps, Networks, and Views
Abstract
Most self-training methods use ensemble averaging across training steps (e.g., exponential moving average (EMA) teachers), networks, or input data transformation views to aid semi-supervised learning. However, little is known about how ensemble size and member weighting affect these averages and thus self-training efficacy. Here, we use a prediction error decomposition to show that variance reduction by training step, network, and view averaging overlaps and is bounded by measurable error correlation. We further identify the private and shared variance of each averaging type by inverting measured correlation, and show gains in effective sample size are sublinearly additive across averaging types. To connect variance reduction to increased task performance, we isolate the contributions of effective sample size gains from pseudo-label quality improvements. By decoupling these components, we show that averaging over larger ensembles approaches an upper bound in shared variance reduction and thus provides quickly diminishing task performance returns, suggesting that targeted reductions in private variance could more efficiently improve self-training. Together with extensive experimentation on Pascal VOC and Cityscapes, we apply our findings to optimize the ensemble averaging configurations of state-of-the-art self-training methods, which improves task performance on the ADE20K, COCO, Pancreas-CT, and Left Atrium datasets. This work establishes error correlation as a common budget that unifies step, network, and view averaging in self-training, and will guide the design of averaging configurations in future methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.