Inverse RL Helps Align AI by Imitating Humans
Abstract
Language model alignment aims to make model behavior reflect desirable properties such as helpfulness, safety, and instruction following. Existing methods draw supervision from expert demonstrations, preference judgments, or verifiable outcomes. Demonstrations typically guide supervised fine-tuning (SFT), while preference judgments and verifiers provide rewards for reinforcement learning (RL). This raises a question: can the demonstrations alone be used to derive an explicit, interpretable reward useful in a post-training pipeline? Motivated by inverse reinforcement learning, we introduce Projected Alignment Reward Estimated from Demonstrations (PARED). PARED learns an explicit reward from expert demonstrations over a small set of response-level features, using a lightweight discriminator that separates demonstrations from the policy's own samples in this feature space. Experiments with inference-time reranking and adversarial on-policy RL show improvements both when starting from a base model without an SFT stage and when further optimizing an SFT-trained model. Additionally, we demonstrate that PARED can be used for contextual alignment, in which a single policy learns to adapt its responses to the preferences of different audiences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.