acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Scalable Planning Paradigms for High-Dimensional Control

Abstract

Test-time planning with learned world models has become a common practice in model-based reinforcement learning in recent years, yet mixed empirical results and a principled limitation arising from actor–planner divergence have also been reported. We revisit whether test-time planning should be a default component of high-dimensional control and re-examine policy improvement through background planning on imagined rollouts. By combining on-policy background planning with an annealed curiosity reward, our method outperforms the strongest prior high-dimensional control method under the matched environment interaction budget, improving mean performance by 8.57% on DMC and 9.99% on HumanoidBench. We further show that curiosity-driven exploration yields larger gains in high-dimensional and sparse-reward settings than either test-time planning or planner-action imitation. In contrast, the benefits of test-time planning and action imitation depend strongly on the task and temperature hyperparameters and are not consistently positive. We conduct an extensive hyperparameter sweep over planner and imitation temperatures. Across the 84 task–configuration cells involving a test-time planner, 21 and 15 yield relative improvements above 10% and relative degradations below -10%, respectively. Based on a correlational analysis, we formulate a causal hypothesis consistent with these observations: early intervention by the test-time planner, when the world model and value function remain underfitted, may degrade action quality and reduce experience diversity in the replay buffer. Together, these results suggest that, under our unified setting, on-policy background planning with curiosity-driven exploration provides a more reliable route to high-dimensional policy improvement than test-time planning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.