acceptodds
Under review as a conference paper at ICLR 2027

Finite-Reference Error Bounds for Recurrent Value Evaluators in Policy Improvement

Abstract

We bound value error for finite-depth recurrent evaluators in policy improvement without assuming recurrent convergence. We compare the endpoint , the value read at the deployed depth , with a deeper finite endpoint of the same evaluator. For a fixed policy, the standard Bellman-residual bound and triangle inequality give an error radius of plus the -step residual of divided by . It is tighter than the direct residual bound on exactly when the residual reduction exceeds ; depth alone does not suffice. When the latent is carried across environment steps, the state includes the latent and the remaining edit budget; fixing the deployed policy and the shared depth- recurrent map keeps both endpoints in one MDP, so changes only endpoint evaluation. Exactly centered advantages connect the radius to a conservative policy improvement lower bound with signed evaluation error. An exact three-state example gives a nonvacuous radius and certifies improvement over the base. On a fixed 128-state Sudoku census, finite-reference proxies are 0.9 to 1.9% smaller than direct proxies across eight training seeds sharing one base initialization. Both remain uninformative as radii at roughly 70 times the checker scale; this is neither a uniform certificate nor a task-performance result. We analyze one candidate mixed exactly, state by state, with a frozen base policy at a fixed parameter snapshot.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.