acceptodds
Under review as a conference paper at ICLR 2027

Scalable Continuous-Time Reinforcement Learning for High-Dimensional Control via Soft Persistence

Abstract

Lots of real-world physical systems evolve continuously over time and require control across diverse timescales. Standard reinforcement learning typically represents continuous control as instantaneous state-to-action decisions, leaving the temporal structure of actions implicit. Continuous-time approaches introduce persistence through adaptive decision timing. However, committing the entire action vector delays feedback-based revisions, which become restrictive in high-dimensional systems with heterogeneous actuator requirements. We analyze the trade-off between commitment loss and value-estimation error amplification, and show how heterogeneous temporal requirements can restrict shared commitment as action dimensionality increases. Our derivation shows that bounded action revision can extend the feasible control horizon beyond hard holding. Motivated by the analysis, we propose Soft Persistence Regularization (SoPer), which learns action dynamics by optimizing task return under a conditional-drift budget. Discretizing this objective yields a practical regularizer that integrates directly into standard actor-critic learning, encouraging temporal coherence while allowing action components to evolve at different rates. Experiments on diverse high-dimensional continuous control benchmarks show sample efficiency advantages of SoPer over discrete-time and continuous-time RL baselines. It can support distinct temporal behavior across actuators in different task and regularization strength. We also demonstrate the success of SoPer in a real humanoid robot experiment. The results suggest soft persistence as a scalable way to incorporating continuous-time action dynamics into efficient learning of control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.