TD-MTP: Coupling Structured Planning with Multi-Step Value Learning
Abstract
In model-based robot learning, a planner selects actions and collects the data used to train its value function. Its effect during learning may therefore depend on the value target. We study this dependence with TD-MTP, which combines a structured tensor-graph planner with replay-based multi-step value targets in TD-MPC2. The main comparison evaluates both planners under one-step and multi-step targets on ten continuous-control tasks. The structured planner improves mean return on one task under one-step targets and five under -returns; its effect is more favorable under -returns on eight. On h1-stair, the planner changes from reducing mean return by to increasing it by . Used separately, the multi-step target raises mean return above the baseline on eight tasks and the planner on one. Controls suggest that elite fitting contributes to this interaction, while the contribution of tensor proposals remains unresolved. For fixed policies, our error bound separates bootstrap perturbation from behavior-prior mismatch and accounts for the endpoint bootstrap retained by finite replay windows. Across 38 task-protocol pairs in four benchmarks, TD-MTP records 10 wins, 22 ties, and 6 losses against TD-MPC2 under stated tolerances. Learning curves show earlier gains on humanoid-run but slower learning on stick-pull. These results support evaluating planners and value targets together during robot learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.