acceptodds
Under review as a conference paper at ICLR 2027

DeCoR: Decoupled Collaborative Reasoning via Preference-Guided Pruning and Dual-Indicator Gating

Abstract

Large Language Models (LLMs) with long Chain-of-Thought (CoT) training demonstrate exceptional reasoning capabilities, yet reliance on excessive reasoning tokens degrades both performance and efficiency. Existing methods address this through specialized retraining or external constraints. While effective, these methods still retain a single-stream reasoning paradigm, leaving exploration and termination entangled within the same trajectory. To overcome this limitation, we propose DeCoR (coupled llaborative easoning), a novel framework for efficient reasoning that decouples exploration and termination through twin-model interaction. We first devise a preference-guided method to achieve structured layer pruning, which derives a functional auxiliary model that serves as an explorer under high predictive uncertainty. This design disentangles distinct reasoning roles while preserving a shared state and trajectory coherence. To coordinate these roles according to the evolving reasoning state, we further introduce a dual-indicator gating mechanism that monitors uncertainty and consistency of intermediate answers. It dynamically alternates between the base model for correction or termination and the auxiliary model for continuous exploration. Extensive experiments on five reasoning benchmarks demonstrate that our method significantly reduces CoT length by 32.5% on average and improves final accuracy by 2.8 to 6.5 percentage points. Our method also increases the density of effective reasoning information within concise outputs. The code will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.