What Can Bad Data Identify? Negative Demonstrations in Offline Decision Making
Abstract
Offline logs often contain crashes, violations, rejected outputs, and abandoned episodes even when expert demonstrations and dense rewards are unavailable. Can these "bad" trajectories reveal what an agent should do, rather than only what it should avoid? Within a finite-MDP, linear-reward framework, we separate three requirements for using failure frequencies to guide control: an observation model that connects latent return to the selection of failures, variation that reveals reward contrasts between competing policies, and transition coverage for the policy ultimately selected. Labels without such an observation model generally identify only actions or trajectories to avoid. Under a negative-rational exponential-family selection model, we characterize reward ambiguity exactly by the null space of centered trajectory features and give an exact test for whether one policy is optimal throughout an affine reward-ambiguity set. The same geometry determines when failures add local Fisher information to an existing log, yields a finite-sample policy-recovery bound under exact-model assumptions, and supports a locally E-optimal design for collecting complementary failure sources. Under matched exponential channels, neither positive nor negative demonstrations are intrinsically more informative: the channel with greater feature variance along the contrasts needed for control has greater local Fisher information, with an exponential advantage possible in either direction. Extensions to negative-unlabeled mixtures and occupancy repulsion expose the additional assumptions behind common uses of bad data. Controlled synthetic tabular experiments isolate the predicted rank, alignment, and coverage effects and demonstrate offline hazard control using failures as the only reward supervision in that setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.