A Contractive Bellman Update Can Solve the Wrong Bellman Equation
Abstract
A contractive Bellman update can have a unique fixed point that recommends the wrong action. We study product eligibility traces where one outcome determines both a temporal-difference error and the coefficient scaling the inherited trace. In finite discounted models with deterministic state-action rewards, the fixed point combines an environmental first transition with an induced continuation kernel. A model can preserve every action gap on its own, yet produce a strict policy cycle through this mixed use. For fixed, reward- and policy-independent continuation kernels, a finite linear test characterizes preservation of every action comparison across rewards and stationary policies. The test permits state-dependent value errors and yields a closest compatible coefficient with unchanged conditional means. Without knowing the transition matrix, we also construct policy-dependent coefficients that can retain outcome dependence while preserving one selected comparison. The induced gap is a strictly positive multiple of the true gap, allowing exact pairwise policy improvement after re-evaluation. A learned-likelihood study connects the failure mechanism to an implemented update: more independent fitting data can make the wrong population ordering more likely. Values need not be exact to support correct policy comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.