Where to Speculate: Cross-Request Survival-Value Allocation for Batched Speculative Decoding
Abstract
Many batched speculative decoders apply a common proposal width to requests whose candidates have sharply different survival probabilities. PACE (Prefix-Aware Compute Allocation) estimates prefix survival before proposal generation, allocates a shared candidate budget across requests, and materializes request-specific widths without batch-wide padding. The verifier remains unchanged. For equal-cost, prefix-consistent proposals, PACE optimizes the supplied survival objective under a fixed budget. The fused runtime is evaluated as an empirical realization of this contract. On the controlled heterogeneous workload, PACE supports – more concurrent requests than the best-static capacity envelope at per-request rate floors of – tok/s. Across a seven-point sweep that varies within-batch survival separation, pre-generation estimates improve equal-budget progress by – over the offline per-budget best-fixed envelope. In the fused I-DLM-8B runtime, materializing request-specific widths within the same N8 container improves throughput by – at concurrency –, showing the effect of omitting unselected candidate positions. Offline replay on external proposal stacks confirms that the allocation relationship transfers beyond the fused decoder; these replay results do not measure physical work removal. Serving gains emerge when the saved candidate computation is large enough to amortize ragged-execution overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.