acceptodds
Under review as a conference paper at ICLR 2027

Learning Adaptive Reasoning Threads for Mathematical Problem Solving

Abstract

Long chain-of-thought (CoT) has substantially improved mathematical reasoning by giving models more computation to develop and refine their solutions. Yet complex problem solving is not naturally linear: some intermediate steps may require extensive reasoning, while later steps need only their conclusions. Linear CoT keeps all such computation in a single growing context, where obsolete details, checks, and failed attempts can interfere with subsequent reasoning. We therefore ask whether a single LLM can learn to selectively develop intermediate derivations separately and reincorporate their conclusions when needed. We introduce Adaptive Reasoning Threads (ART), an on-demand, state-driven protocol that lets the model selectively extend reasoning from its evolving main-thread state into private local threads. Each extension inherits the problem and reasoning accumulated so far, continues with the same model policy, and returns only its conclusion when needed. We learn this behavior through supervised fine-tuning followed by outcome-based reinforcement learning. Across model scales and competition-mathematics benchmarks, ART consistently outperforms matched linear-CoT baselines, including longer-CoT controls with comparable total generation budgets. Simply exposing the ART interface is insufficient to reproduce these gains, showing that effective extension behavior must itself be learned. Moreover, ART-trained models remain stronger under standard CoT decoding and exhibit substantially fewer corrections and restarts. These results suggest that learning when to extend local reasoning can improve mathematical problem solving beyond simply increasing reasoning length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.