Think Simpler to Solve Harder: Difficulty-Aware Adaptive Reasoning via Token Entropy
Abstract
Decomposing complex problems into subproblems can improve the reasoning capabilities of large language models (LLMs). However, choosing an appropriate degree of decomposition based on subproblem difficulty and limiting the impact of intermediate-answer errors on final results remain important challenges. We find that answer-token entropy, defined as the mean entropy of token distributions across positions in a generated answer, can reflect subproblem difficulty at suitable granularities. Building on this observation, we propose **DART**, a training-free algorithm for **D**ifficulty-aware **A**daptive **R**easoning through **T**oken entropy. At suitable granularities identified through one-time probing, DART uses answer-token entropy as a routing signal to adaptively choose between direct solving and further decomposition, balancing solution quality and inference overhead. We also design a confidence-aligned self-critique mechanism that jointly constrains candidate evaluation scores and score-token confidence to guide the evaluation and refinement of intermediate answers, restraining the impact of errors on subsequent reasoning. Results aggregated across three tasks show that, compared to the baseline with the lowest average error on each backbone, DART reduces average error by 22%–57% and average online inference cost by 51%–72%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.