acceptodds
Under review as a conference paper at ICLR 2027

The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective

Abstract

Weight-perturbation evolution strategies (ES) can fine-tune billion-parameter language models with surprisingly small populations (e.g., ), contradicting classical zeroth-order curse-of-dimensionality intuition. A second, seemingly separate phenomenon also appears under fixed hyperparameters, when the fine-tuning reward often rises, peaks, and then degrades, not just with ES but also with GRPO. Both effects admit a common geometric account: fine-tuning landscapes are *low-dimensional in curvature*, with a few stiff directions embedded in a large, nearly flat bulk. The separation of curvature scales produces (i) heterogeneous time scales that yield rise-then-decay under fixed stochasticity, as captured by a minimal quadratic stochastic-ascent model, while the concentration of curvature produces (ii) degenerate improving updates, where many random perturbations share similar components along the stiff directions. For isotropic perturbations, directions without curvature add nothing to the cost of a random perturbation, so the probability of sampling an improving one does not depend on the ambient dimension , and the scaling question reduces to how the curvature-active dimension grows with . Probing the fine-tuning reward landscapes of GSM8K, ARC-C, and WinoGrande with ES across Qwen2.5-Instruct models (0.5B-7B) reveals that reward-improving perturbations remain accessible with small populations at every scale, and direct Hessian-spectrum measurements on the same models yield effective curvature dimensions on the order of tens, growing sub-linearly with over the tested scales and occupying a vanishing fraction of parameter space. These results reconcile ES scalability with non-monotonic training dynamics under a single geometric account, and they open the door to a broad class of fine-tuning algorithms, beyond policy-gradient RL, that worst-case zeroth-order theory had ruled out as infeasible at scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.