The Executable Content of Gibbs Path Tilts
Abstract
Gibbs trajectory reweighting can favor environmental outcomes that a controller cannot choose. Yet an infeasible path target need not induce suboptimal actions. We quantify this distinction in finite-horizon entropy-regularized control. Building on established information projections and residual-value correction, we express the discrepancy as an auxiliary control problem whose cost is local transition optimism. Its propagated value gives an exact criterion for lossless action extraction and an action-variance identity. As the cost scale tends to zero, relaxation optimism is and execution loss is ; either leading coefficient may vanish. We derive explicit coefficients over the full Markov decision process and characterize a sixth-order degenerate case. Under learned dynamics, small model error can reverse the benefit of correction near a decision tie. A confidence-certified selector protects the regularized objective; a paired transfer bound exploits cancellation between candidates. Supplied aggregate records cover 663,552 random control problems and 12,779,520 finite-sample trials. In the latter, unconditional correction is harmful in 38.46% of trials; the evaluated conservative selector switches in 7.06%, with no recorded harmful switch on a valid confidence event. These studies support the gap mechanisms and learned-model comparison, but do not test the asymptotic rates or paired refinement. The contribution is a quantitative account of decision-relevant path optimism, not a new control solver.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.