HintMR: Eliciting Stronger Mathematical Reasoning in Small Language Models
Abstract
Small language models (SLMs) often struggle with complex mathematical reasoning, where errors made early in a solution can compound over long reasoning trajectories. We introduce HintMR, a hint-assisted reasoning framework that elicits stronger reasoning in SLMs by providing localized guidance throughout multi-step problem solving. At each reasoning step, a separate language model generates a context-aware hint conditioned on the problem and the solver's accumulated reasoning history. Rather than revealing the full solution, these hints provide targeted intermediate guidance that helps the SLM stay on a productive reasoning trajectory while continuing to generate the solution itself. We further introduce Adaptive HintMR (AdaHintMR), an efficient variant that selectively invokes the hint model only when additional guidance is likely to be useful. AdaHintMR estimates the solver's step-level uncertainty using token-level Shannon entropy and requests a new hint only when uncertainty exceeds a predefined threshold, allowing the SLM to reason independently on confident steps. Experiments across six mathematical reasoning benchmarks show that HintMR consistently improves reasoning accuracy over standard prompting and independent SLM reasoning. When paired with GPT-5.4 as the frontier hinter, HintMR substantially improves the reasoning performance of SLM solvers across multiple benchmarks, outperforming GPT-5.4 used directly as a standalone solver by up to 6.67 percentage points. These gains show that using a frontier model for targeted intermediate guidance can be more effective than relying on it to solve the entire problem directly. AdaHintMR further reduces reliance on the frontier model by selectively requesting hints only when the solver exhibits high uncertainty. Across the evaluated solvers and benchmarks, this reduces API cost by up to 72.4% relative to HintMR while maintaining strong reasoning performance. Overall, our results show that targeted intermediate hints can effectively elicit stronger mathematical reasoning from SLMs, while uncertainty aware hinting preserves much of this benefit with substantially lower frontier model inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.