Zeroth-Order Momentum and Adaptive Methods for Nonsmooth Policy Optimization from Human Preferences
Abstract
Reinforcement learning from human feedback (RLHF) commonly relies on reward models to convert human preferences into optimization signals, but reward inference introduces an additional modeling step and may suffer from reward misspecification. To avoid reward inference, direct policy optimization methods learn policies directly from human preferences without explicitly fitting a reward model, for which zeroth-order optimization provides a natural optimization framework. Existing zeroth-order policy methods such as zhang2025zeroth rely on smooth policy objectives, leaving the nonsmooth setting less understood. Meanwhile, momentum has shown favorable theoretical and empirical properties, motivating its study in preference-based zeroth-order policy optimization. In this work, we study zeroth-order policy optimization from human preferences for general objectives that may be nonsmooth and nonconvex. We establish a refined characterization of the preference oracle induced by pairwise human feedback for reward-model-free RLHF. We further establish a unified convergence framework for momentum-based and adaptive-gradient zeroth-order nonsmooth policy optimization methods, called ZOPGM and ZOPGAM. Both methods achieve a -expected stationarity (a natural relaxation of -Goldstein stationarity) outer-loop complexity of . This rate matches the parameter dependence obtained by translating the best-known zeroth-order Goldstein-stationarity bound to expected stationarity. Accounting for preference-oracle sampling yields a total human preference query complexity of . Experiments in controlled preference-based policy optimization environments demonstrate consistent improvements in preference-query efficiency over zeroth-order baselines across different preference-feedback budgets, supporting the view that momentum helps exploit information accumulated from historical preference-induced updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.