acceptodds
Under review as a conference paper at ICLR 2027

Where to Pause and How to Resume: Towards Human-Inspired Iterative Chain-of-Thought

Abstract

Iterative chain-of-thought (CoT) reasoning extends reasoning depth by periodically summarizing intermediate steps and resuming generation within a bounded context. However, existing iterative CoT approaches typically construct supervision through length-based segmentation, which can interrupt reasoning at arbitrary positions and provide limited guidance on how to resume from summaries. We propose PhaseCoT, a data-centric refinement that improves iterative reasoning along two dimensions: where to pause and how to resume. PhaseCoT segments reasoning trajectories at semantic milestones, so that each segment corresponds to a coherent reasoning phase, and further rewrites segment heads to introduce context-aware continuation from prior summaries. We construct supervised fine-tuning datasets for standard CoT, conventional length-driven iterative CoT, and PhaseCoT, and evaluate multiple models on MATH500, AIME22, and AMC23. Results show that PhaseCoT improves mathematical reasoning over both standard CoT and length-driven iterative CoT in most model-benchmark comparisons, which suggests that effective iterative reasoning requires not only longer reasoning horizons, but also semantically grounded pauses and context-aware resumptions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.