CoTDog: Closed-Loop Control for Efficient LLM Reasoning
Abstract
Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve increasingly complex problems by allocating more computation to intermediate reasoning. However, this additional computation is not always used effectively, as reasoning models may spend excessive tokens on intermediate steps or repeatedly revisit already explored directions. In this paper, we introduce CoTDog, a closed-loop framework that guides reasoning progression and answer production within a single generation process. An external controller monitors visible trace events and resource use while maintaining persistent execution state. It guides the model to control its reasoning pace, move to the next stage when needed, and produce the final answer at the right time. We evaluate CoTDog on three open-weight reasoning models (i.e., QwQ-32B, Qwen3.5-27B, and DeepSeek-R1-Distill-Qwen-32B) across MATH-500, GPQA Diamond, and LiveCodeBench. Across the nine settings,CoTDog achieves 7.9% higher task scores and uses 23.7% fewer completion tokens, averaging the relative differences against the strongest baseline in each setting. On Qwen3.5-27B with LiveCodeBench, it improves pass@1 by 18.8 percentage points over the strongest baseline while using 22.1% fewer tokens. The results show that appropriate guidance during CoT reasoning can substantially improve reasoning performance without requiring model parameter updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.