acceptodds
Under review as a conference paper at ICLR 2027

Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning

Abstract

Many structured reasoning problems can be formulated as iterative updates to an explicit state under global constraints. Recent recursive models show that this form of computation can be effective in task-specific networks, but it remains unclear whether pretrained language models can acquire the same capability, how intermediate training states should be selected, and whether such adaptation trans- fers beyond the source task. We study these questions by coupling a pretrained language-model backbone with a tied recurrent updater over explicit answer states. We introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects intermediate states from model-generated trajectories according to the finite- horizon progress achieved by the current model. Rather than selecting states only by their source or initial difficulty, FOCUS prioritizes states according to how effectively they can be improved by the learned transition. Across all evaluated language-model backbones, recurrent state updates consistently improve structured- task performance over direct and one-step training. On Maze-Hard, for example, Qwen3-1.7B reaches 86.0% accuracy with Recurrent final-only, compared with 0.0% for both Direct LoRA SFT and One-step State-SFT. Beyond the source tasks, the resulting language-model adapters exhibit zero-shot transfer to general reasoning benchmarks: with the recurrent module disabled at evaluation and no downstream parameter updates, Sudoku-trained Qwen3-1.7B improves MATH- Hard from 72.36% to 72.96% and CruxEval-O from 50.75% to 55.38%. These results show that pretrained language models can acquire strong explicit-state repair capability and that structured repair training can generalize beyond the recurrent solver, transferring to mathematical and code reasoning under zero-shot evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.