Recur, Switch, Recur, Commit: Function-Adaptive Hybrid Reasoning
Abstract
Large language models can reason through explicit tokens or recurrent latent states, but hybrid systems still lack a principled rule for deciding which representation to use at each step. We show that this choice is governed more by local reasoning function than by position: latent recurrence better supports search and planning by retaining alternative candidates, whereas explicit computation better supports calculation and verification by preserving exact values and entities. Controlled interventions reverse these preferences as candidate multiplicity and exact-value demand change, and frozen hidden states predict them before router training. We therefore introduce Function-Adaptive Hybrid Reasoning, combining capability-preserving multi-path training, function-grounded switching supervision, and budgeted router-only reinforcement learning for bidirectional, state-conditioned routing. Across eight model configurations and twelve core tasks, the full method improves the twelve-task macro over explicit CoT-RL on every backbone, by 0.8–2.2 percentage points. On Qwen3-8B, it gains 1.9 points while using 0.72 normalized compute and a 0.75 estimated normalized latency. These results suggest that effective hybrid reasoning is less about when to compress thought than about matching each local computation to the representation that best supports it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.