FOCUS: Scaling Continuous Control via Structured Exploration over Diversified Action Subspaces
Abstract
Reinforcement learning in high-dimensional continuous action spaces faces a significant bottleneck: the exponentially expanding search space leads to poor sample efficiency and limited exploration scalability. To address this challenge, we propose FOCUS, a framework designed for structured exploration that avoids the exhaustive search of the full action space. The core insight of FOCUS lies in decoupling control into a hierarchical selection of diversified action subspaces, effectively rendering the exploration process more tractable. Specifically, a high-level policy generates binary preferences to dynamically activate task-critical control units, while a low-level policy produces precise actions conditioned on these selections. To ensure the effectiveness of this hierarchy, we incorporate a diversity-consistency tradeoff in the optimization objective. This objective encourages functional specialization by promoting distinct action distributions between preferences that select a control unit and those that do not, while encouraging consistency across preferences that leave the unit unselected. Extensive evaluations on challenging continuous control benchmarks demonstrate that FOCUS significantly outperforms state-of-the-art methods, particularly as the action dimensionality increases, establishing its superior scalability and exploration efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.