acceptodds
Under review as a conference paper at ICLR 2027

Displacement-Aware Subspace Adaptation for Muon-Style Zeroth-Order Fine-Tuning

Abstract

First-order fine-tuning of large language models (LLMs) is often constrained by the substantial memory overhead of backpropagation, whereas zeroth-order (ZO) fine-tuning offers a memory-efficient alternative based solely on forward evaluations. However, high-dimensional ZO estimates can be noisy, making it difficult to reliably exploit matrix structure through Muon-style spectral updates. Recent ZO-Muon methods alleviate this issue by restricting ZO estimation to low-rank search subspaces, but repeatedly refreshing these subspaces at random may discard directions that have consistently driven recent parameter updates. Consequently, subsequent Muon steps may operate in less informative search spaces, as polar normalization can only exploit directions already captured by the selected bases. To address this issue, we propose Displacement-Aware Subspace ZO-Muon (DAZM), which uses accumulated parameter displacements produced by recent Muon updates to adapt the low-rank search subspace. Specifically, we accumulate the realized Muon updates over each interval, allowing directions that are repeatedly reinforced along the optimization trajectory to become prominent in the resulting displacement. We then retain the leading left and right singular directions as trajectory-informed exploitation bases, while filling the remaining dimensions with random directions orthogonal to the retained subspaces to preserve exploration. Theoretically, we show that under the temporal persistence of the local optimization geometry, displacement-aware subspace adaptation retain more informative gradient structure for subsequent Muon updates than purely random refreshes. Extensive experiments on multiple LLMs and downstream tasks demonstrate that our method achieves consistently superior performance over existing ZO-Muon and non-Muon ZO baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.