Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
Abstract
Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offer dramatic cost savings yet lag substantially in accuracy. Existing approaches—routing and cascading—treat the LLM as an all-or-nothing resource: either the query bypasses the LLM entirely, or the LLM generates a complete response at full cost. We find that the LLM need not be used this way: even a truncated prefix of its response raises SLM accuracy, with gains that broadly grow with prefix length on all four benchmarks we study, also it hold across model families, tokenizers, and SLM sizes. We therefore introduce LLM Shepherding, which requests only a short prefix (a *hint*) from the LLM and lets the SLM produce its own complete response. Shepherding generalizes both routing and cascading, and it achieves lower cost under oracle decision-making. We develop a two-stage predictor that jointly determines whether a hint is needed and how many tokens to request. On GSM8K, CNK12, HumanEval, and MBPP, we compare shepherding with routing and cascading. Across full cost–accuracy sweeps, Reactive Shepherding attains the largest area under the cost–accuracy curve on all four benchmarks; at a fixed target of 90% of LLM accuracy, it is the cheapest method on all four; and at default operating points it reduces cost by 33–65% relative to LLM-only inference and delivers up to 4x cost reduction compared to state-of-the-art routing and cascading baselines. To our knowledge, shepherding is the first to formalize and learn a single hand-off prefix hint, controlled only through `max_new_tokens`, under a black-box per-token cost model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.