acceptodds
Under review as a conference paper at ICLR 2027

Beyond Selection: Recycling Noisy Real-World Preference Data for Post-Training

Abstract

Preference data based on real-user queries capture how people use language models in various tasks and contexts. Many examples, however, have limited value for post-training in their original form. Data-quality pipelines often approach this with selection: retain useful examples and discard the rest. We study a complementary question: can discarded examples instead be recycled to provide more useful supervision? Our framework, ReSource, leverages noisy queries as sources for new tasks, allowing the task to change while keeping meaningful connections to the underlying content. We train a source-grounded instruction generator and use the quality of induced responses as a proxy for instruction utility to guide both training and instruction selection. In comparisons of recycled supervision against raw supervision using the same source examples, we find that recycling consistently improves downstream performance across three benchmarks and different amounts of training data. We further investigate its value when mixed with retained data and when used to train other models. Overall, our findings suggest that the utility of an example for post-training in its original form does not fully determine its value as a source of supervision, motivating recycling as a complement to data selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.