acceptodds
Under review as a conference paper at ICLR 2027

MetaThink: Learning When to Think via Multi-Token Metacognitive Reasoning

Abstract

Large reasoning models achieve strong performance through chain-of-thought (CoT) reasoning, but uniformly applying lengthy deliberation to all problems incurs substantial computational overhead. Existing adaptive reasoning methods train models to selectively engage CoT, yet some rely on expensive and inconsistent external annotations, and they typically use a single-token decision mechanism that creates an information bottleneck and leads to decision boundary collapse. Inspired by metacognitive gating in human dual-process theory, we propose MetaThink, which introduces a multi-token metacognitive reasoning stage where the model generates a brief natural language self-assessment before deciding whether to engage CoT. Combined with an asymmetric reward that assigns higher rewards to correct non-CoT responses, MetaThink enables the model to autonomously discover which problems require reasoning through reinforcement learning, without external difficulty annotations. Experiments on the mathematical reasoning domain show that MetaThink achieves a competitive accuracy-efficiency tradeoff compared to existing adaptive reasoning methods across multiple model scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.