acceptodds
Under review as a conference paper at ICLR 2027

Stiefel Manifold Optimization Accelerates Reinforcement Learning

Abstract

Reinforcement learning with high-dimensional observations and overparameterized neural networks is common in practice, yet its learning dynamics remain poorly understood. To develop a theoretically grounded approach to improving learning in this setting, we study learning a feedback controller for the linear-quadratic regulator problem with high-dimensional observations generated from low-dimensional latent states. We use deep linear networks to model overparameterized policies and value functions, with weight matrices constrained to have orthonormal columns, i.e., to lie on the Stiefel manifold. Within this framework, we show that the learning dynamics of the high-dimensional optimization problem match those of the underlying low-dimensional latent-state problem, allowing actor-critic learning to exploit the hidden low-dimensional structure. Motivated by this theory, we improve the performance of proximal policy optimization by incorporating Stiefel-manifold optimization in MuJoCo environments with pixel observations, trained end-to-end from reward and without auxiliary representation-learning losses. We further show that the same optimization principle improves large language model fine-tuning using distillation for the Qwen3-4B model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.