acceptodds
Under review as a conference paper at ICLR 2027

ONE VOCABULARY PASS, MANY REQUESTS: A STREAM-ONCE OUTPUT-HEAD KERNEL FOR CONCURRENT LLM DECODING

Abstract

The output head of a large-vocabulary language model is a maximum-inner-product search. Existing exact CPU implementations batch queries or reuse blocked products but do not make a single compact vocabulary-row stream the shared screening unit across concurrent decoder states. FuseHead performs one affine-u8 row traversal, screens each row–query pair, and completes the survivor union in FP32. A public 24-query panel and two held-out groupings verify all canonical FP32 row scores and final token decisions, with additional cancellation, one-ULP, extreme-scale, random, and tie-breaking controls. On an eight-core CPU and a Qwen2.5-0.5B-derived WikiText-2 trace, FuseHead matches all 288 FP32 decisions and reduces median output-head time from 2.542 seconds to 0.591 seconds. Its paired headline speedup over an exact per-panel oracle is 2.03× [1.18,4.10]. When measured no-cache prefix replay is included, end-to-end throughput is 1.09× that of the strongest same-information dense comparator.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.