acceptodds
Under review as a conference paper at ICLR 2027

A Primer in Post-Training Reasoning Data: What We Know About How It Works

Abstract

Post-training has become a major driver of recent progress in large reasoning models, and reasoning data are often a key determinant of its success. Research on these data has grown rapidly, yet the evidence remains fragmented. This primer synthesizes over 150 works around four questions: what forms reasoning data take, what makes them useful, how they are constructed, and what scaling studies establish. Across mathematics, code, science, and interactive tasks, we compare apparently conflicting findings to clarify the conditions under which conclusions hold and the questions that remain open. We distill this evidence into seven practical lessons and reporting recommendations to guide data construction, assess feedback reliability, and evaluate claims of data efficiency and capability improvement. We conclude by outlining directions for future research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.