When to Think Deeply: Response-Conditioned Inhibitory Routing for LLM Reasoning
Abstract
Selective reasoning in language models must balance opportunities to correct an available answer against additional computation and harmful replacements. Inspired by inhibitory control in the human brain, we introduce Inhibitory Deliberation for Problem Reasoning (IDPR), which separates answer formation from answer release. A fast policy generates a candidate, and a response-conditioned gate either releases it or inhibits its release to invoke slow reasoning. Decomposed supervision uses paired fast–slow outcomes to learn slow-over-fast quality gain, generation cost, and corrective potential from the question, fast answer, and generation evidence. Calibration selects the most accurate of three scoring rules at fixed 8% coverage. On a fixed-policy benchmark of 5,000 mathematical problems, IDPR raises accuracy from 47.90% to 48.92% at an 8.20% slow-call rate, with 519.42 average generated tokens. The gain over Always-Fast is significant under a Holm-adjusted paired McNemar test (), with 111 corrected errors and 60 harmful replacements. With response-and-evidence inputs and encoder family held fixed, IDPR-Matched's joint supervision-and-scoring scheme achieves 1.40 percentage points higher accuracy than RouteLLM-Q+R+E while reducing slow calls from 9.46% to 7.14% at their respective validation-calibrated thresholds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.