Structured Reasoning for Large Language Models
Abstract
Large reasoning models (LRMs) achieve strong performance by producing long chains of thought (CoT). However, a single trajectory entangles functionally different reasoning behaviors, making it difficult to assess or optimize each behavior on its own merits. This entanglement is reinforced by prevailing training recipes, where supervised fine-tuning imitates all tokens with uniform weight and outcome-level reinforcement learning compresses the entire trajectory into a single outcome reward, leaving the model without ability-specific credit and unable to optimize individual behaviors. To address this, we propose StruCtured Reasoning (SCR), which reorganizes reasoning into stages that are explicit, evaluable, and independently trainable, and realizes them through a Generate-Verify-Revise paradigm. Specifically, SCR introduces two trajectory-aware supervision mechanisms, where Dynamic Termination Supervision ties the end of reasoning to the self-verification outcome, and Selective Loss Masking excludes incorrect initial answers from the imitation loss while preserving them as context. A progressive two-stage reinforcement learning schedule then jointly optimizes generation and verification, and turns to revision once the verifier becomes sufficiently reliable. Across three backbone models and multiple reasoning benchmarks, SCR consistently improves reasoning quality and self-verification accuracy while reducing output length by up to 50%. Code is available at https://anonymous.4open.science/r/SCR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.