What Action Geometry Buys: A Design Space for Optimal-Transport Bellman Backups
Abstract
Many Reinforcement Learning (RL) action spaces carry geometry that maximum-entropy backups ignore. Laplacian optimal transport (LOT) regularization admits a closed form—an average of cost-shifted local softmax—governed by one anchor-normalized action kernel; we turn this closed form into a design space, proving which kernel property buys which guarantee. Concentration buys planning: exact pruning error and a localized planner replace branching over actions by an effective count , still targeting the unpruned regularized value. Uniform cost stability buys learning: reusable reward-free geometry and same-stream tabular Q-learning. Poisson truncation buys local computation: a graph-local kernel targets the exact heat-kernel value with no explicit dependence under polynomial growth. Strictly positive couplings buy interpretation: inverse policy curvature is effective resistance on a value-dependent transport graph, reducing to Fisher–Rao at zero cost. Matched obstructions show these properties are not free; choosing a geometry is choosing a row of our map.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.