LAMP: Lyapunov-Inspired Potentials for Latent Model Predictive Control
Abstract
Latent model predictive control ranks candidate actions using a learned model and a bootstrapped value function. During learning, errors in these components can distort candidate preferences, particularly when experience is limited. We introduce **LAMP**, a planning extension that learns an auxiliary state potential from reward-bin labels and uses it to penalize candidate trajectories. The potential uses a Lyapunov-inspired non-increase objective on sampled reward-improving transitions. It is trained on detached latents without Bellman-return targets. It guides planning through a scheduled hinge penalty. Our analysis characterizes the penalty's effects on candidate weights, conditions for reducing absolute and relative score errors, and limits of transferring reward supervision to planning queries. On 28 robotic control tasks from the DeepMind Control Suite, the reported results show improved learning performance over TD-MPC2, with the largest gains on Dog and Humanoid locomotion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.