Policy Improvement as Decision Revision: Preserving the Incumbent under Executable Feedback
Abstract
Executable counterfactual evaluation provides evidence about alternative actions, but utilities alone do not specify the decision that this evidence is meant to revise. We formalize the evaluation object as an evaluated decision, (p_old, Q), pairing the incumbent action distribution with executable utilities over a common semantic-action support. Projecting this object onto Q is many-to-one: any supervision constructed solely from Q is invariant to the incumbent and therefore discards its identity before policy improvement begins. This motivates a representation principle for learning from executable feedback: revise the evaluated decision relative to its incumbent rather than replace it with a decision constructed from utilities alone. Distributional Decision Revision (DiRe) instantiates this principle using an established KL-regularized improvement target and an incumbent-relative residual policy parameterization. We test this principle through controlled interventions that separately isolate revision specification and finite-model realization. With executable utilities fixed, changing the incumbent leaves Q-only targets unchanged but changes the DiRe revision. On matched evaluated decisions, Hard-Q, utility-only supervision, and DiRe encode similar expected executable improvements (+0.0645, +0.0625, and +0.0638) while inducing substantially different incumbent-relative displacements (JS 0.1101, 0.1554, and 0.0688), showing that executable gain alone does not characterize the revision. We then hold the DiRe target, model family, cross-entropy objective, optimization budget, and initialization fixed and intervene only on the realization reference. On 59 state-held-out decisions, using the evaluated incumbent rather than a uniform reference reduces target-policy JS from 0.1313 to 0.0572 and changes realized executable improvement from -0.0162 to +0.0053. Exact reconstruction and richer-scorer analyses further separate specification from realization: the target is representable in principle, while a semantic residual scorer realizes 50.9% of its intended improvement on a later state-disjoint batch. Together, these results identify incumbent identity as information that can be lost at the evaluation-to-supervision interface, while finite-model realization remains a distinct and heterogeneous bottleneck. Independent conversation-level generalization remains untested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.