When Does More Information Pay Off? Effective Cost in LLM Inference
Abstract
More context is not always cheaper. Extra information can reduce how much an LLM needs to generate, but a longer prompt also makes each cached decoding step more expensive. We study when the saved computation outweighs the cost of carrying additional context. We define effective cost as the smallest physical budget needed to reach a common target, such as a desired response behavior or task score. Under regularity conditions, one representation provides at least as much attainable utility as another across budgets and decision objectives exactly when it reaches every response behavior attainable from the other at no greater effective cost. We then derive when a larger context is worthwhile and develop rules for jointly choosing the input context and the amount of generation, including the cost of constructing or replacing the cached state. In controlled routing, useful extra information lowers the cost of reaching a common numerical success threshold despite a longer prompt. On 576 held-out MuSiQue questions with supplied intermediate answers, jointly choosing how many paragraphs to retain and how long to generate improves mean answer F1 by 0.040 over choosing one evidence scope per budget, averaged across three prespecified profiled GPU-time budgets. Finally, rebuilding an active KV cache from a shorter prompt yields decoding savings that exceed the measured rebuilding cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.