acceptodds
Under review as a conference paper at ICLR 2027

Learning from Failures: When Contrastive Post-Training Beats Success Conditioning

Abstract

Post-training a language model from rollouts with binary outcomes has two standard recipes: supervised fine-tuning on the successful rollouts (success conditioning), and pairwise contrastive learning on success–failure pairs. In practice the two are combined, and the combination is often reported to work best, but there is little theory of when and why. We analyze the combined family: a margin-based contrastive term plus a weighted success conditioning term on successful rollouts using a contextual bandit setup with binary rewards. We characterize when adding the contrastive term improves on success conditioning, both in value and other metrics such as those used in retrieval (e.g., precision, recall). We also propose another combined family that contains a DPO-style contrastive loss term that accounts for the base policy. We compare it with the combined loss that contains the more traditionally used triplet loss under different settings. We validate our theoretical results via synthetic experiments and also in the real-world context of generative retrieval.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.