acceptodds
Under review as a conference paper at ICLR 2027

Capacity-Aware Policy Improvement for Model-Based Reinforcement Learning

Abstract

Model-based reinforcement learning (MBRL) can combine a learned policy with online planning for data-efficient continuous control. However, in planner-guided learning, the planner collects experience, whereas actor updates evaluate actions sampled from the nominal policy. This policy-planner mismatch can expose actor updates to inaccurate value estimates for actions poorly covered by replay data. Critic-based weighting can prioritize promising actions, but concentrating weight on a few samples can amplify the influence of errors in their value estimates. We propose Capacity-Aware Policy Improvement (CAPI), which formulates weighting within a finite actor batch as an entropy-regularized allocation problem with an explicit capacity constraint. The resulting allocation reweights the TD-MPC2 actor objective, prioritizing high-scoring actions while limiting the normalized weight share assigned to each sample. Experiments on high-dimensional continuous-control tasks from the DeepMind Control Suite and HumanoidBench demonstrate that CAPI improves data efficiency and final performance compared to strong baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.