acceptodds
Under review as a conference paper at ICLR 2027

How Does Loss Landscape Geometry Transfer Across Scales?

Abstract

Neural scaling limits, such as the maximal update parameterization (P), have become a powerful tool for zero-shot hyperparameter transfer as model size grows. But which aspects of the loss landscape itself exhibit analogous transfer across scales? We study this question through the lens of *linear mode connectivity* (LMC), which asks whether models are connected in weight space by a linear path of low loss. We find evidence that interpolation barriers approach scale-consistent behavior, can already be stable at practically relevant widths and depths, and that such transfer can persist over long training horizons, beyond the regime of local quadratic approximation. Experiments with convolutional networks and transformers for vision and language show consistent behavior across width and depth under a range of training settings, and demonstrate that this consistency extends to practical weight-space operations such as model merging. As a theoretical foundation for our observations, we establish a scaling limit for loss barriers in a simplified setting. Our analysis offers a unifying perspective on neural scaling theory and weight space geometry, laying a foundation for understanding when interpolation, weight averaging, and model merging behave consistently across scales.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.