acceptodds
Under review as a conference paper at ICLR 2027

Where to Speculate: Cross-Request Survival-Value Allocation for Batched Speculative Decoding

Abstract

Many batched speculative decoders apply a common proposal width to requests whose candidates have sharply different survival probabilities. PACE (Prefix-Aware Compute Allocation) estimates prefix survival before proposal generation, allocates a shared candidate budget across requests, and materializes request-specific widths without batch-wide padding. The verifier remains unchanged. For equal-cost, prefix-consistent proposals, PACE optimizes the supplied survival objective under a fixed budget. The fused runtime is evaluated as an empirical realization of this contract. On the controlled heterogeneous workload, PACE supports – more concurrent requests than the best-static capacity envelope at per-request rate floors of – tok/s. Across a seven-point sweep that varies within-batch survival separation, pre-generation estimates improve equal-budget progress by – over the offline per-budget best-fixed envelope. In the fused I-DLM-8B runtime, materializing request-specific widths within the same N8 container improves throughput by – at concurrency –, showing the effect of omitting unselected candidate positions. Offline replay on external proposal stacks confirms that the allocation relationship transfers beyond the fused decoder; these replay results do not measure physical work removal. Serving gains emerge when the saved candidate computation is large enough to amortize ragged-execution overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.