acceptodds
Under review as a conference paper at ICLR 2027

Learning from Graded Feedback: Tier-Aware Self-Play Listwise Optimization

Abstract

Offline preference optimization relies on fixed response comparisons that become increasingly mismatched to an evolving policy. Self-play mitigates this mismatch by regenerating responses from the current policy. Extending self-play to listwise learning is challenging because multi-constraint evaluations often assign multiple responses the same discrete score, producing ordered tiers with ties. Strict-order listwise objectives require tie-breaking, which introduces unsupported preferences, while binary pass/fail feedback provides no relative signal when all rollouts are imperfect. We propose Tier-Aware Self-Play Listwise Optimization (T-SPLO), a framework for learning from graded verifier feedback. At each round, prompts are prioritized according to their current tier gap, and the policy generates multiple rollouts that an objective verifier groups into tiers according to constraint satisfaction. T-SPLO optimizes the likelihood of the resulting tiered order by marginalizing over all within-tier permutations using subset dynamic programming. The resulting objective is invariant to permutations among equally scored responses and learns from different degrees of partial success. Across two backbones and three benchmarks, T-SPLO achieves the highest final average instruction-following score among the compared baselines, with end-to-end ablations favoring tie preservation and average scores improving across training rounds. Our code is available at https://anonymous.4open.science/r/T-SPLO-956D/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.