acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Taylor Attribution for Sequential Policy Learning

Abstract

Existing reinforcement-learning approaches have shown strong performance in sequential active perception by optimizing acquisition policies from long-term utility. However, these methods typically model only the next acquisition conditioned on the current observation history, leaving alternative multi-step acquisition strategies implicit in the policy. We introduce a diffusion policy with Counterfactual Taylor Attribution (CTA) to explicitly exploit such future plan variation. The diffusion policy models a joint distribution over multi-step acquisitions, enabling diverse candidate plans to be sampled and counterfactually evaluated from the same decision state. CTA uses the resulting utility variation to estimate action-level contributions and redistribute within-plan credit during policy optimization, without introducing an additional reward model or changing the original task objective. Our framework provides a general mechanism for exploiting counterfactual plan variation in active information acquisition, including active perception and sequential Bayesian experimental design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.