Stability and Decision Fidelity in Tree-Structured Bellman Updates
Abstract
Tree-structured sequential protocols aggregate child continuations before local Bellman updates. Normalizing this channel can reduce its sensitivity while changing the decision problem. Tree-SeBIL distinguishes stability of a declared semantic operator from preservation of its decisions, separating fixed-point-preserving averaging from target-changing surrogates. On a closed invariant domain, explicit reward, transition, and policy-recovery bounds yield a structured small-gain certificate; a restricted leaf-injection family gives a sharp unweighted norm threshold. On an action-rich class, universal local Gibbs-policy fidelity at fixed continuation inputs and unchanged rewards holds exactly for aggregation translations. Reducing the aggregation Lipschitz budget below the semantic aggregation's minimal modulus yields a worst-case continuation mismatch and binary policy-separation bound, sharp on the stated linear family. An independent 900-map interception study illustrates low surrogate residual with substantial original-objective decision change. Neural results provide limited, mixed wrapper evidence, not equilibrium certificates. Fixed points solve the declared induced-node problems; transfer to a game's regularized equilibrium additionally requires policy consistency and compatible reward ownership, transitions, and histories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.