Logic Signatures for Test-Time LLM Reasoning Routing
Abstract
Reasoning in Large Language Models (LLMs) is commonly summarized by benchmark scores and improved with a uniform inference recipe. However, this scalar view neglects the interaction between a query’s logical demands and a particular solver’s failure tendencies. We hypothesize a failure-to-control principle: once a solver’s failures are indexed by the reasoning operations that elicit them, its error history can be inverted into a query–solver-specific policy for selecting assistance before the same vulnerabilities recur. To test this hypothesis, we propose LOGICROUTE, a training-free framework for pre-response intervention routing. At its core is LOGICSIG, a sparse two-sided representation that separates what the query requires from where the solver tends to fail. Using offline zero-shot traces, LOGICROUTE constructs a solver-specific logic profile, organizes candidate interventions into general reasoning scaffolds and targeted repairs, and routes each query before the target model responds. Across five Qwen models spanning 1.7B–27B parameters on MATH500, AIME, GPQA, and MMLU-Pro, LOGICROUTE improves macro accuracy from 51.00% to 62.51%, outperforming semantic demonstration retrieval by 8.19 percentage points. Across the evaluated Qwen, Llama, and Mistral models, macro accuracy increases from 39.55% to 51.79%. Profile substitutions show that useful routes are solver-specific; route perturbations support the utility of logic-feature ordering; and demonstration-bank ablations show that assistance is compositional but not monotonic in context volume. Together, these findings suggest that failures in the past provide a reusable pre-response control signal for targeted assistance without parameter updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.