Scaffold Builder: Diagnosis-Guided Data Synthesis from Intervenable Failures
Abstract
The bottleneck of post-training is shifting from insufficient data scale to the scarcity of high-learning-value data. Existing data synthesis methods are often decoupled from the model's current capability state: the generated samples may be more complex or higher quality, yet they do not necessarily target the model's actual capability gaps. To address this issue, we propose Scaffold Builder, a diagnosis-guided data synthesis framework that formulates data construction as a teaching process centered on the model's capability boundary. Inspired by educational scaffolding and the Zone of Proximal Development (ZPD), Scaffold Builder organizes synthesis into an iterative loop of boundary mining, scaffold construction, controlled synthesis, quality filtering, and feedback training. It first identifies samples near the student model's capability boundary, after which a teacher model diagnoses the corresponding failures into a structured RootCause–Fix–Scaffold triplet. The effectiveness of each scaffold is then verified through student-model re-evaluation. Guided by the validated scaffolds and diagnosed root causes, the framework synthesizes targeted training data along two complementary paths and uses a three-stage curriculum to shift from scaffolded learning toward independent answering. Experiments across multiple reasoning benchmarks show that, with only 300 seed examples, Scaffold Builder achieves higher average performance than strong data synthesis and distillation baselines. Further ablation, iterative, model-scale, training-paradigm, and cross-domain experiments demonstrate the effectiveness, scalability, and transferability of the proposed framework.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.