acceptodds
Under review as a conference paper at ICLR 2027

Clipped Optimization over Hierarchical Networks under Generalized Smoothness

Abstract

Hierarchical learning systems combine federated aggregation within groups of clients with decentralized communication across hubs. We study optimization in this setting under -smoothness, where the local smoothness level may grow with function suboptimality. This model is particularly convenient for hierarchical objectives because averaging preserves the curvature-growth parameter , with heterogeneity entering through an additive interpolation defect. The main difficulty is locality, since the corresponding descent inequality is valid only for sufficiently small displacements. We therefore use clipping at both levels of the hierarchy: clients clip gradients before aggregation, while hubs clip the optimization step and leave the gradient tracker unchanged. We prove nonconvex convergence guarantees for the resulting federated, decentralized, and hierarchical methods, with accelerated dependence on the network spectral gap. For convex objectives, the same mechanism leads to two phases: linear convergence while clipping is active, followed by an regime. Under strong convexity, both phases are linear.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.