acceptodds
Under review as a conference paper at ICLR 2027

A Contractive Bellman Update Can Solve the Wrong Bellman Equation

Abstract

A contractive Bellman update can have a unique fixed point that recommends the wrong action. We study product eligibility traces where one outcome determines both a temporal-difference error and the coefficient scaling the inherited trace. In finite discounted models with deterministic state-action rewards, the fixed point combines an environmental first transition with an induced continuation kernel. A model can preserve every action gap on its own, yet produce a strict policy cycle through this mixed use. For fixed, reward- and policy-independent continuation kernels, a finite linear test characterizes preservation of every action comparison across rewards and stationary policies. The test permits state-dependent value errors and yields a closest compatible coefficient with unchanged conditional means. Without knowing the transition matrix, we also construct policy-dependent coefficients that can retain outcome dependence while preserving one selected comparison. The induced gap is a strictly positive multiple of the true gap, allowing exact pairwise policy improvement after re-evaluation. A learned-likelihood study connects the failure mechanism to an implemented update: more independent fitting data can make the wrong population ordering more likely. Values need not be exact to support correct policy comparisons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.