acceptodds
Under review as a conference paper at ICLR 2027

How Does Atomic Data Composition Shape LLM Mathematical Reasoning?

Abstract

Reinforcement learning (RL) has substantially advanced the mathematical reasoning capabilities of Large Language Models (LLMs). However, how the domain composition of training data shapes target-capability amplification remains unclear. Motivated by the field-atomic perspective, we investigate how supervised fine-tuning (SFT) data from algebra, geometry, analysis, and topology affects subsequent target-domain GRPO. Beyond final performance, we annotate reasoning chains using ThinkARM's eight reasoning behaviors and analyze behavior distributions, transition dynamics, and their conditional associations with correctness. Our results show that domain-aligned SFT does not consistently provide the strongest initialization, largest GRPO gain, and best final performance, with the most effective source domain varying across target capabilities. Different target domains exhibit distinct behavioral demands, while models initialized with different field data retain different reasoning profiles after the same target-domain GRPO. Moreover, the strength and direction of transition–correctness associations vary across **domains capability** and **training compositions** and Path-reward experiments show that reinforcing high-value transitions can improve target-domain reasoning. These findings highlight the importance of behavior-aware data selection and post-training for more effective mathematical capability amplification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.