A Primer in Post-Training Reasoning Data: What We Know About How It Works
Abstract
Post-training has become a major driver of recent progress in large reasoning models, and reasoning data are often a key determinant of its success. Research on these data has grown rapidly, yet the evidence remains fragmented. This primer synthesizes over 150 works around four questions: what forms reasoning data take, what makes them useful, how they are constructed, and what scaling studies establish. Across mathematics, code, science, and interactive tasks, we compare apparently conflicting findings to clarify the conditions under which conclusions hold and the questions that remain open. We distill this evidence into seven practical lessons and reporting recommendations to guide data construction, assess feedback reliability, and evaluate claims of data efficiency and capability improvement. We conclude by outlining directions for future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.