acceptodds
Under review as a conference paper at ICLR 2027

CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment

Abstract

Reasoning-capable language models often produce long chains of thought when direct answers suffice, wasting inference compute. Dual-mode models offer both thinking and direct-answer modes, but typically leave mode selection to users. Automating this selection while improving responses under both modes is challenging because routing targets evolve with the policy, strong initial mode preferences can destabilize exploration, and sequence-level objectives entangle routing with response learning. We introduce CounterRoute, an online reinforcement-learning framework that jointly learns routing and mode-conditioned responses in one shared policy directly from a native dual-mode checkpoint, without method-specific SFT warm-up. Paired current-policy counterfactual rollouts estimate cross-mode routing credit, which is assigned only to the routing token, while within-mode GRPO trains response tokens. A paired-to-self-routed curriculum stabilizes early training with forced coverage of both modes before progressively transferring training to autonomous routing. Across nine benchmarks, CounterRoute better balances accuracy and efficiency than heuristic and learned adaptive-routing methods, improving macro-average accuracy at both Qwen3 scales while cutting mean generated tokens by 51% and 41%, respectively, versus always-thinking checkpoints. On instruction-following and commonsense benchmarks where direct answering is strong, think rates fall as low as 1% while response quality improves. Despite training only on math and synthetic instruction-following data, its routing and response quality generalize to held-out tasks in coding, science, knowledge, and commonsense.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.