acceptodds
Under review as a conference paper at ICLR 2027

Target Identification in Offline Imitation Learning: The Effect of State Conditioning

Abstract

We study reward-free offline imitation learning with logs from a preferred and a comparison behavior source. Their differences can be used to score candidate behavior. The resulting target is the state-action visitation distribution that maximizes this score while remaining feasible under the environment's dynamics and close to the logged behavior. We ask whether source visitation frequencies alone determine that target when the audit does not use linked transitions. A joint score uses both where the sources go and what actions they choose there; conditioning on state removes the first kind of information. We show that this removal can make the target ambiguous: the same two source occupancies uniquely determine every joint target in a two-state example, yet permit different conditional targets under different compatible dynamics. Under positivity and when the sources do not span all feasible flow directions, we prove that a Kullback-Leibler (KL) regularized target is uniquely determined exactly when its unconstrained reweighting is a nonnegative occupancy in the affine span of the sources. For finite samples, State-Conditional Occupancy Reweighting (SCOR) returns a target estimate and a confidence ball covering every target compatible with the true source occupancies, without estimating transitions. Sampling uncertainty shrinks with more data for identified targets with bounded affine coefficients, while uncertainty caused by indistinguishable dynamics can persist. Experiments across 30 source geometries examine these claims and the value of signed affine combinations. The guarantees concern the target specified by the learning objective, not task return.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.