Learning Soft Constraints from Reversed and Recoverable Preference Pairs
Abstract
Soft constraints shape how a response should be written, asking for a writing style, a situation to respect, or a scope of content. Unlike hard constraints, they cannot be verified by code, so existing pipelines build preference data for them by asking an LLM judge which response is better. Whether a response follows a soft constraint is a subjective call, so the judge's verdict carries its own biases, and a soft constraint does not always change the response enough to give the judge anything to compare. We present REVERB, whose principle is to plant, between the two responses of a pair, a difference that is easy to tell apart and to learn from. Reversing one soft constraint into its contrary gives two mirrored prompts, and each response is labeled as preferred under the prompt that produced it. Whether the two responses really separate along the reversed constraint is then verified: an LLM sees only the two responses and must recover the reversed constraint, a test with a known answer, and pairs that fail it are discarded. A single open-weight model runs the whole pipeline and produces 7,949 verified pairs. Direct preference optimization (DPO) on this data improves constraint following on two benchmarks by up to 16.4 points on two student models, outperforming every compared baseline, beating the judge-labeled pipeline by up to 12.0 points with fewer pairs, and preserving general abilities across five benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.