Xtra-MTP: State-Efficient Parallel Verification for Hybrid Language Models
Abstract
Native multi-token prediction (MTP) uses a model’s own prediction module to draft future tokens for joint target verification. In hybrid language models, which combine full-attention and recurrent layers, this opportunity faces three coupled constraints: uneven numbers of accepted tokens, duplicated candidate states, and dependencies that limit verification parallelism. We present Xtra-MTP, which jointly designs candidate coverage, recoverable state, and Gated DeltaNet execution. One-token alternatives extend one recursively drafted path without recursively expanding additional branches. A shared checkpoint and compact update records preserve candidate continuations; the next verifier recovers only the selected path, sharing state access with verification. A topology-aware solver separates updates that propagate along the path from terminal readouts, while value tiling exposes parallel execution. Across three Qwen models and six workloads, with the 95th percentile of request-average time per output token (TPOT) constrained to 30 ms, measured configurations improve six-workload geometric-mean throughput by 29.4–39.2% over the best tested SGLang configurations and 5.9–9.2% over our ReplaySSM integration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.