acceptodds
Under review as a conference paper at ICLR 2027

When to Wait Before Fan-Out: Warm-Up, Verification, and Residency in Cache-Aware LLM Evaluation

Abstract

Long-context LLM evaluation often issues many requests that share a long prefix. When the provider does not coordinate cold requests, dispatching them concurrently can create the same prefix repeatedly, whereas serial release sacrifices parallelism. We study this release decision from the client side, where cache state is hidden and reuse is observed only through completed-response telemetry. Cache-Aware Execution Planning (CAEP) changes only execution order and release time, keeping request payloads and scoring fixed; its stable-canary policy (CAEP-SC) releases a prefix family after a leader and a probe verify reuse, and a break-even rule relates avoided prefix creation to the value of added latency. In development studies, CAEP-SC reduces telemetry-priced cost relative to work-conserving scheduling by 49.9% on DeepSeek at a 3.7% makespan increase, and creation-confirmed release reduces it by 63.4% on Alibaba's explicit cache at 1.50× makespan. Matched controls locate these savings in warm-up rather than verification: in an eight-block hosted ablation, leader-only and blind two-request warm-up match the gate's cost in every block, and across 37 replicated matched comparisons gating never resolves a cost or computation saving. Under constrained vLLM capacity, CAEP-SC processes 3.18× the work-conserving input with Llama-3.1-8B, 50.7% more than blind warm-up; Qwen3.6-27B shows a 41.2% gate penalty at an 8 GiB budget, and competing prefixes can erode reuse after a passing probe. Limiting unfinished families removes excess synthetic computation; on a fixed Qwen QA workload, a four-family cap cuts computed input by 51.2% and makespan by 32.1% relative to uncapped blind warm-up, although work-conserving release remains faster and the effect is unresolved on new document banks. Waiting before fan-out pays off when concurrent cold requests are billed separately, and one completed warm-up request captured the full measured hosted saving.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.