acceptodds
Under review as a conference paper at ICLR 2027

Global Optimization via Softmin Energy Minimization

Abstract

Non-convex optimization in high dimensions requires balancing gradient information with global exploration, a tension that Langevin dynamics and other gradient-informed methods address only partially. We introduce a gradient-based swarm method driven by a Softmin energy interaction , a smooth approximation of the minimum over the ensemble that couples particles through their relative energies, and add Brownian noise to its gradient flow. We show that at every stable equilibrium the interaction splits the ensemble into particles at local minima and particles at local maxima; that without noise the ensemble's best value converges at least as fast as gradient flow under a Polyak–Lojasiewicz condition; and that the interaction caps the Freidlin–Wentzell quasipotential barrier between adjacent basins at order , independently of the true barrier height and strictly below the barrier an overdamped Langevin particle faces. Consistently with this cap, on double-well and quadruple-well benchmarks Softmin reaches the target in every run at every barrier height tested, while Simulated Annealing, Langevin dynamics, and variants of Particle Swarm Optimization and Consensus-Based Optimization all eventually fail; counting failed runs, it is more than an order of magnitude cheaper than Simulated Annealing overall, though not on low barriers. Nor is this a mere effect of ensemble size: Softmin succeeds in every run with five particles, whereas Simulated Annealing gains from more particles only what independent restarts predict, enough on moderate barriers but not on high ones. On Lennard-Jones clusters, Softmin's energy gap to the best known minimum stays nearly flat as the cluster grows while every baseline's rises, and the method trains transformers with millions of parameters. Its limits are equally clear: Particle Swarm Optimization is faster on a high-dimensional Ackley function, and Adam on transformer training, where no method struggled.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.