Lifted Bellman Linear Programming for Offline Reinforcement Learning
Abstract
Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets, which are stabilized by target networks updated through exponential moving averages (EMA). Furthermore, multi-step targets incorporate behavior-policy actions and therefore require off-policy correction. We instead impose in-sample Bellman optimality on the critic through inequality constraints. We formulate the Lifted Bellman Linear Program (LBLP), which lifts the linear programming characterization of Bellman optimality to the joint space so that every constraint involves only state-action pairs in the dataset. Its unique minimizer is the in-sample optimal pair, and constraints along -step segments of dataset trajectories leave this minimizer unchanged for any rollout policy and any horizon. Under deterministic dynamics, this minimizer lies between the best return of the dataset and the optimal value. Relaxing the constraints into hinge penalties recovers the same solution above a finite penalty coefficient in the tabular case. Approximate Lifted Bellman Unconstrained Minimization (ALBUM) implements this relaxation with neural networks and detaches the -step rollout targets by stop gradient. Its objective contains no squared regression onto bootstrapped targets, so it trains without target networks or EMA updates. Under deterministic dynamics, the -step constraints allow discounted returns along dataset trajectories to serve as lower bounds without off-policy correction or action chunking, providing a direct mechanism for long-horizon value propagation. On OGBench, ALBUM uses a single critic with a Gaussian policy, matches the average performance of FQL, and substantially outperforms ReBRAC, while using the fewest parameters and the least peak GPU memory among all measured methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.