DeCoR: Decoupled Collaborative Reasoning via Preference-Guided Pruning and Dual-Indicator Gating
Abstract
Large Language Models (LLMs) with long Chain-of-Thought (CoT) training demonstrate exceptional reasoning capabilities, yet reliance on excessive reasoning tokens degrades both performance and efficiency. Existing methods address this through specialized retraining or external constraints. While effective, these methods still retain a single-stream reasoning paradigm, leaving exploration and termination entangled within the same trajectory. To overcome this limitation, we propose DeCoR (coupled llaborative easoning), a novel framework for efficient reasoning that decouples exploration and termination through twin-model interaction. We first devise a preference-guided method to achieve structured layer pruning, which derives a functional auxiliary model that serves as an explorer under high predictive uncertainty. This design disentangles distinct reasoning roles while preserving a shared state and trajectory coherence. To coordinate these roles according to the evolving reasoning state, we further introduce a dual-indicator gating mechanism that monitors uncertainty and consistency of intermediate answers. It dynamically alternates between the base model for correction or termination and the auxiliary model for continuous exploration. Extensive experiments on five reasoning benchmarks demonstrate that our method significantly reduces CoT length by 32.5% on average and improves final accuracy by 2.8 to 6.5 percentage points. Our method also increases the density of effective reasoning information within concise outputs. The code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.