acceptodds
Under review as a conference paper at ICLR 2027

Enhancing Model Metacognition via Cognitive Pairwise Training

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become central to LLM reasoning, but its outcome-level rewards can make models more willing to give confident answers when evidence or reasoning is insufficient. Existing SFT or RL methods mainly teach LLMs to refuse or express uncertainty at the response level, which can overfit abstention behavior rather than improve reasoning reliability. To address this limitation, we propose **Cognitive Pairwise Training (CPT)**, a cognitive mid-training alignment stage that directly trains the policy to compare the relative quality of reasoning traces, with the goal of learning a reasoning-quality signal that remains useful in subsequent task optimization. Across five model scales and three model families, CPT improves the **reasoning–metacognition trade-off** and substantially enhances the **robustness of subsequent RL**, yielding a – smaller decline in abstention ability than standard SFT and eliminating the decline at 14B. At 14B, CPT+RL outperforms the standard SFT+RL pipeline by math-average points and abstention-F1 points. Further analyses show that CPT improves trace quality and exhibits strong robustness and scalability across evaluation and training settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.