Think Forward, Eliminate Backward: Training and Routing Complementary Reasoning Strategies in LLMs
Abstract
Large language models are trained almost exclusively on forward chain-of-thought (CoT): reasoning proceeds from premises toward a conclusion. Yet a second natural mode of human reasoning — elimination — proceeds by generating candidate answers, evaluating each against the problem's constraints, ruling out the incorrect ones, and converging on the survivor. Prior work has shown that LLMs prompted to reason by elimination perform poorly , but whether elimination can be trained into a model as a learned reasoning strategy, and whether the two strategies are complementary at the problem level, remains underexplored. We fine-tune Qwen3-8B with LoRA on teacher-generated elimination traces (generate candidates evaluate rule out converge) and forward CoT traces, and study the resulting strategy space. We find that (i) elimination reasoning can be successfully trained into an 8B model, (ii) the two strategies exhibit substantial complementarity — 5–29% of problems per dataset are solved by exactly one strategy — and (iii) a lightweight per-query router over TF-IDF and hand-crafted features yields a +2.4% gain over the best single strategy and exceeds the dataset-level oracle, at a modest 13% increase in token cost. Unlike model-level routing , our router selects between reasoning strategies on the same model, adding no deployment cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.