Local M-Covers for Replicated Learning
Abstract
Replicated learning temporarily introduces multiple copies of a model, modifying the training dynamics in ways that can improve generalization while allowing the replicas to contract back to a single deployable model. We ask whether the internal structure of this temporary lift can be tuned so that these replica fluctuations induce task-relevant corrections. We study this through an -cover, which replicates and rewires a model's factor graph, and introduce a local -cover in which replica fluctuations interact through structured rather than uniform coupling. For a general model, we derive the dynamics of replica-difference modes near synchronization and show that generalization improves when their induced predictor displacement aligns with the source-posterior correction. In an analytically tractable random-feature model with two trainable linear layers, this yields a concrete link between source (teacher-prior)–representation mismatch and the preferred cover geometry, and predicts the transition between improved and degraded generalization. Across broader benchmarks, structured -covers outperform established replicated-learning methods across tasks, architectures, learning dynamics, and overparameterization regimes. These results identify the internal geometry of temporary replication as a new design axis for generalization: structured replica interactions can improve the final single-model predictor without increasing deployment-time model size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.