LEGO-MATH: Stacking Reasoning Blocks for Reliable and Harder Math Problem Synthesis
Abstract
High-quality mathematical reasoning data is essential for improving the reasoning capabilities of large language models, but obtaining such data at scale remains challenging. Existing LLM-based synthesis methods either rely on free-form prompting, which often produces invalid problems or incorrect solutions, or use rigid templates, which improve controllability but limit the diversity and difficulty of generated problems. In this work, we propose LEGO-MATH, a compositional framework for synthesizing mathematical reasoning problems by representing solutions as directed acyclic graphs (DAGs) and recombining them through compatible reasoning interfaces. LEGO-MATH first converts existing problem-solution pairs into reasoning DAGs, where nodes correspond to intermediate reasoning states and edges encode inferential dependencies. It then identifies compatible interfaces via edge-aware node clustering and iteratively stacks compatible DAG fragments, starting from an initial pair and extending the current composite until an adaptive stopping condition is met. Since the logical structure of the synthesized problem is inherited from verified source DAGs, correctness can be preserved by propagating only localized numerical updates along the affected downstream subgraph after each composition, reducing the role of the LLM to simple recomputation rather than open-ended generation. This structure-preserving composition enables LEGO-MATH to jointly improve validity, covering both problem and solution correctness, and effectiveness, reflected by increased problem difficulty. Extensive experiments show that LEGO-MATH produces more valid and more challenging mathematical reasoning data, and that the resulting data provides useful training signal for post-training large language models. Our code and examples of synthesized data are available at https://anonymous.4open.science/r/LEGO-MATH-PUBLIC-5D0E.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.