LIFT: Scoped Constraints for Evaluating and Training Instruction Following
Abstract
Common instruction following (IF) benchmarks, such as IFEval, evaluate the ability of language models (LMs) to follow simple constraints that can easily be verified programmatically. These same constraints, in turn, are used as reward signals for reinforcement learning. This design may lead to reward hacking of constraints, pushing models away from the main task. Recent attempts to address this resulted in benchmarks with richer constraints, but required expert labor, remained closed-access, and cannot be easily extended. We present LIFT-Bench, an IF benchmark that asks a model to draft a document given a topic and a structured outline. Constraints are scoped to certain document sections, providing finer control over constraint density. LIFT-Bench includes both verifiable and judged constraints, as well as a training counterpart, LIFT-Train. Importantly, LIFT-Bench provides both main task and instruction-following constraints, mitigating shallow reward hacking. On experiments with multiple models, we find that frontier models struggle with the harder subsets of our benchmark and analyze model performance along multiple axes. We also find that training on LIFT-Train is much less susceptible to reward hacking than training on the common IFTrain. LIFT-Bench and LIFT-Train are publicly available, as well as the data generating pipeline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.