acceptodds
Under review as a conference paper at ICLR 2027

TRIBE: Trust-Region Policy Improvement from Batch Experience

Abstract

Post-training language models from scalar rewards is commonly approached with policy-optimization methods developed for more general sequential decision problems. In sequence-level post-training, treating a complete response as a single action allows policy improvement to be written directly in terms of how the candidate policy reweights sampled responses. This motivates , a two-stage method that first solves for the desired policy improvement in importance-ratio space and then updates the language model to realize it. The resulting formulation provides a unified approach to on-policy, off-policy, and fully offline training, including settings where the policies that generated the offline data are unavailable, and it allows additional policy constraints to be incorporated directly into the improvement step. Across GSM8K and MATH, TRIBE is competitive with strong on-policy baselines and is the strongest behavior-policy-independent method in our controlled offline experiments. On UltraFeedback, TRIBE is the only evaluated method to improve over rejection fine-tuning on both win rate and reward advantage. We also show that an expected-cost constraint can be added without changing the model-fitting stage, yielding competitive safety-constrained alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.