acceptodds
Under review as a conference paper at ICLR 2027

Imitation Beyond Expectation via Second-order Stochastic Dominance

Abstract

Expected utility theory provides an ordering for policies when the utilities of behavior sequences are given. Many imitation learning methods extend this perspective by seeking policies that match or exceed the utility of demonstrations in expectation, despite the exact utility function only being assumed to be Markovian (and linear). However, this reduces a distribution of demonstrations to its average characteristics, ignoring the distinct characteristics of individual demonstrations that may be important to imitate. We adopt second-order multivariate stochastic dominance, a stricter notion for partially ordering distributions, to guide imitation learning. It broadens the performance guarantees of imitation learning to all non-decreasing and concave utility functions. We investigate both left-sided (i.e., risk-avoidant) and right-sided (i.e., risk-seeking) variants of this stochastic dominance to demonstrate the advantages compared to employing stricter first-order stochastic dominance and other imitation learning baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.