acceptodds
Under review as a conference paper at ICLR 2027

Predicting Block-Coordinate Performance via Cross-Curvature

Abstract

Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss–Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an remainder, where is the learning rate. We derive a signed loss comparison after iterations with error under regularity conditions and for fixed , identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0% of iterations for the neural network, 83.4% for federated learning, and 97.6% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0%, 93.2%, and 99.6%, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.