Learning to Switch with Stability and Convergence Guarantees: Policy Gradient for State-Independent Controller Switching in Linear Systems
Abstract
The design of feedback controllers plays a central role in determining how dynamical systems behave under closed-loop operation. A single fixed controller, however, may not provide the desired behavior for every performance objective. Addressing different requirements by repeatedly redesigning the controller can be difficult, especially when the design space is continuous. When a finite set of pre-designed controllers is available, the problem can be reduced to deciding how these controllers should be selected over the operating time horizon, motivating controller switching. However, the key challenge of switching is to improve closed-loop performance while maintaining stability. This is especially important in online learning, where an unstable update can destabilize the system during training. To address this, we develop StableSwitch-PG, a policy- gradient framework that optimizes a state-independent controller-selection distribution while preserving mean-square stability throughout learning for linear time- invariant systems. For known dynamics, we derive the exact policy gradient and establish global convergence of the objective to the optimal randomized switching policy. For unknown dynamics, we estimate the gradient from sampled trajectories and establish asymptotic and finite-time convergence guarantees. Finally, we derive a Lyapunov-based update condition that preserves mean-square stability at every learning iteration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.