Understanding and Selecting Reasoning Strategies in LLM Agents
Abstract
We study how different reasoning paradigms affect language model performance and when each is useful. Comparing direct answering, chain-of-thought, planning, reflection, and tool use reveals that strategy rankings vary across models and tasks, while multiple strategies can solve the same question at different computational costs. These observations motivate select-then-solve (STS), which learns strategy values from execution outcomes and token costs, then chooses a strategy before solving a new question. STS retains both successful alternatives and failed attempts as supervision. On the original in-distribution split, its cost-aware variant improves score by 2.57 percentage points over a training-selected fixed strategy while using 7.2% fewer solver tokens. Comparisons with dataset lookup and held-out task families indicate that task-family priors account for much of the observed benefit. Controlled experiments reveal instruction–tool interactions and compare repeated sampling within and across configurations. Together, our findings inform when and how to select reasoning strategies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.