Counterfactual Taylor Attribution for Sequential Policy Learning
Abstract
Existing reinforcement-learning approaches have shown strong performance in sequential active perception by optimizing acquisition policies from long-term utility. However, these methods typically model only the next acquisition conditioned on the current observation history, leaving alternative multi-step acquisition strategies implicit in the policy. We introduce a diffusion policy with Counterfactual Taylor Attribution (CTA) to explicitly exploit such future plan variation. The diffusion policy models a joint distribution over multi-step acquisitions, enabling diverse candidate plans to be sampled and counterfactually evaluated from the same decision state. CTA uses the resulting utility variation to estimate action-level contributions and redistribute within-plan credit during policy optimization, without introducing an additional reward model or changing the original task objective. Our framework provides a general mechanism for exploiting counterfactual plan variation in active information acquisition, including active perception and sequential Bayesian experimental design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.