Step Dependency Awareness in LLMs: Measuring and Improving Structural Understanding via Permutation-Based Training
Abstract
Chain-of-thought (CoT) reasoning improves LLM performance, yet it remains unclear whether models genuinely understand structural dependencies among reasoning steps or merely pattern-match surface correlations. We formalize step dependency awareness—the ability to distinguish order-constrained (dependent) from order-free (independent) steps—and introduce the Permutation Robustness Probe, which compares model behavior under valid (dependency-preserving) and invalid (dependency-violating) step permutations via a Dependency Awareness Score (DAS). Across six models (0.5B–8B; general, base, math, and code variants), models exhibit implicit sensitivity to ordering (DAS_conf up to 5.60) yet fail to detect invalid orderings explicitly (accuracy 50%). To narrow this gap we propose Dependency-Aware Permutation Training (DAPT), which applies Direct Preference Optimization (DPO) to valid vs. invalid permutation pairs sharing the same final answer, isolating ordering from answer correctness. On Qwen2.5-1.5B-Instruct, DAPT raises GSM8K accuracy from 51.5% to 64.7% 1.6% (+13.2%, 3 seeds) while achieving the first above-chance detection discrimination (87.7% valid acceptance, 22.2% 4.1% invalid rejection, up from 12.3%) and a 3 step-ordering gain on the primary seed (2.5% to 7.5%) that rises to 4 on the three-seed mean (2.5% to 10.0% 2.0%). In contrast, SFT on original chains collapses to answer extraction (DAS = 0) and SFT on randomly permuted chains degrades both accuracy (19%) and sensitivity (DAS_conf = 0.15). These results show that structural understanding can be measured and improved through ordering-only training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.