acceptodds
Under review as a conference paper at ICLR 2027

When Does More Information Pay Off? Effective Cost in LLM Inference

Abstract

More context is not always cheaper. Extra information can reduce how much an LLM needs to generate, but a longer prompt also makes each cached decoding step more expensive. We study when the saved computation outweighs the cost of carrying additional context. We define effective cost as the smallest physical budget needed to reach a common target, such as a desired response behavior or task score. Under regularity conditions, one representation provides at least as much attainable utility as another across budgets and decision objectives exactly when it reaches every response behavior attainable from the other at no greater effective cost. We then derive when a larger context is worthwhile and develop rules for jointly choosing the input context and the amount of generation, including the cost of constructing or replacing the cached state. In controlled routing, useful extra information lowers the cost of reaching a common numerical success threshold despite a longer prompt. On 576 held-out MuSiQue questions with supplied intermediate answers, jointly choosing how many paragraphs to retain and how long to generate improves mean answer F1 by 0.040 over choosing one evidence scope per budget, averaged across three prespecified profiled GPU-time budgets. Finally, rebuilding an active KV cache from a shorter prompt yields decoding savings that exceed the measured rebuilding cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.