acceptodds
Under review as a conference paper at ICLR 2027

Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference

Abstract

Large Language Models (LLMs) deliver state-of-the-art performance on complex reasoning tasks, but their inference costs limit deployment at scale. Small Language Models (SLMs) offer dramatic cost savings yet lag substantially in accuracy. Existing approaches—routing and cascading—treat the LLM as an all-or-nothing resource: either the query bypasses the LLM entirely, or the LLM generates a complete response at full cost. We find that the LLM need not be used this way: even a truncated prefix of its response raises SLM accuracy, with gains that broadly grow with prefix length on all four benchmarks we study, also it hold across model families, tokenizers, and SLM sizes. We therefore introduce LLM Shepherding, which requests only a short prefix (a *hint*) from the LLM and lets the SLM produce its own complete response. Shepherding generalizes both routing and cascading, and it achieves lower cost under oracle decision-making. We develop a two-stage predictor that jointly determines whether a hint is needed and how many tokens to request. On GSM8K, CNK12, HumanEval, and MBPP, we compare shepherding with routing and cascading. Across full cost–accuracy sweeps, Reactive Shepherding attains the largest area under the cost–accuracy curve on all four benchmarks; at a fixed target of 90% of LLM accuracy, it is the cheapest method on all four; and at default operating points it reduces cost by 33–65% relative to LLM-only inference and delivers up to 4x cost reduction compared to state-of-the-art routing and cascading baselines. To our knowledge, shepherding is the first to formalize and learn a single hand-off prefix hint, controlled only through `max_new_tokens`, under a black-box per-token cost model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.