Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Abstract
Reinforcement learning (RL) research has begun to shift focus toward alignment, ensuring agents learn behaviors that adhere to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches that combine these signals are largely designed for the contextual bandit framing of language generation. Within this framing, recent work has started moving beyond multi-stage pipelines toward fully offline, single-stage pipelines that can reduce computational overhead and remove the need for environment access, broadening alignment's applicability. Yet little work extends such pipelines to sequential decision-making settings, where these complementary inputs can serve as a richer, interconnected signal. We propose Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies. We adapt Safety Gymnasium environments as a principled testbed for alignment evaluation, on which FMR reduces misalignment by up to 98% across a range of imitation learning algorithms while often improving aptitude. FMR remains robust in limited-data regimes, even when demonstrations become predominantly noisy and uninformative.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.