CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment
Abstract
Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Dual-mode models offer both thinking and direct-answer modes, but typically leave mode selection to users. Automating this selection while improving responses under both modes is challenging because routing targets evolve with the policy, strong initial mode preferences can destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and mode-conditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts estimate cross-mode routing credit, which is assigned only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced coverage of both modes before progressively transferring training to autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods, improving macro-average accuracy at both Qwen3 scales while cutting mean generated tokens by 51% and 41%, respectively, versus always-thinking checkpoints. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and synthetic instruction-following data, its routing and response quality generalize to held-out tasks in coding, science, knowledge, and commonsense.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.