acceptodds
Under review as a conference paper at ICLR 2027

Stepwise Think-Critique: Interleaved Reasoning and Self-Critique in a Single LLM

Abstract

Human beings solve complex problems through critical thinking, where reasoning and evaluation are intertwined to converge toward correct solutions. However, most existing large language models (LLMs) treat the reasoning and verification as separate processes: they either generate reasoning without explicit self-checking or rely on external verifiers to detect errors post hoc. The former lacks immediate feedback, while the latter increases system complexity. Motivated by human critical thinking, we propose Stepwise Think-Critique (STC), an end-to-end trainable framework in which a single LLM emits a structured, step-level critique inline with each reasoning step. STC is trained with reinforcement learning that complements reasoning rewards with a critique-consistency reward derived from final-answer correctness, jointly optimizing reasoning correctness and critique reliability. On five mathematical reasoning benchmarks, STC retains the reasoning gains of standard RL (+7.2 Pass@1 points over its 1.5B base model) while substantially improving its step-level self-critique (F1 from 46.4% to 67.4%). In judging its own reasoning steps, it on average outperforms seven dedicated process reward models of up to 8B parameters, without any threshold tuning. STC takes a step toward LLMs with built-in critical thinking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.