acceptodds
Under review as a conference paper at ICLR 2027

Inverse RL Helps Align AI by Imitating Humans

Abstract

Language model alignment aims to make model behavior reflect desirable properties such as helpfulness, safety, and instruction following. Existing methods draw supervision from expert demonstrations, preference judgments, or verifiable outcomes. Demonstrations typically guide supervised fine-tuning (SFT), while preference judgments and verifiers provide rewards for reinforcement learning (RL). This raises a question: can the demonstrations alone be used to derive an explicit, interpretable reward useful in a post-training pipeline? Motivated by inverse reinforcement learning, we introduce Projected Alignment Reward Estimated from Demonstrations (PARED). PARED learns an explicit reward from expert demonstrations over a small set of response-level features, using a lightweight discriminator that separates demonstrations from the policy's own samples in this feature space. Experiments with inference-time reranking and adversarial on-policy RL show improvements both when starting from a base model without an SFT stage and when further optimizing an SFT-trained model. Additionally, we demonstrate that PARED can be used for contextual alignment, in which a single policy learns to adapt its responses to the preferences of different audiences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.