Capacity-Aware Policy Improvement for Model-Based Reinforcement Learning
Abstract
Model-based reinforcement learning (MBRL) can combine a learned policy with online planning for data-efficient continuous control. However, in planner-guided learning, the planner collects experience, whereas actor updates evaluate actions sampled from the nominal policy. This policy-planner mismatch can expose actor updates to inaccurate value estimates for actions poorly covered by replay data. Critic-based weighting can prioritize promising actions, but concentrating weight on a few samples can amplify the influence of errors in their value estimates. We propose Capacity-Aware Policy Improvement (CAPI), which formulates weighting within a finite actor batch as an entropy-regularized allocation problem with an explicit capacity constraint. The resulting allocation reweights the TD-MPC2 actor objective, prioritizing high-scoring actions while limiting the normalized weight share assigned to each sample. Experiments on high-dimensional continuous-control tasks from the DeepMind Control Suite and HumanoidBench demonstrate that CAPI improves data efficiency and final performance compared to strong baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.