acceptodds
Under review as a conference paper at ICLR 2027

Efficient EAGLE-3 Training via Structure-Aware Execution

Abstract

Speculative decoding speeds up large language model inference using a lightweight drafter to propose multiple candidate tokens which the target model checks in parallel. EAGLE-3 is a recent draft-model architecture for speculative decoding. The data used to train drafters have heterogeneous supervision ratios. That is, the fraction of valid token positions with direct supervision varies across samples and datasets. Existing EAGLE-3 training frameworks still execute recurrent computations that cannot affect the training loss and recompute token-side projections for inputs shared by different recurrent paths, incurring avoidable computation and memory costs. We propose structure-aware execution based on two observations: some recurrent updates cannot affect a supervised prediction, and different computations can repeatedly project the same token-side input. These observations lead to three execution changes. First, we keep only recurrent computations needed by selected losses at the current or later training-time test (TTT) steps, while retaining their same-path predecessors and the full public key/value (K/V) context required by the retained queries. Second, we design fused attention to process these potentially non-contiguous queries without storing public-only outputs and log-sum-exp (LSE) values for a later merge. Third, we compute and reuse a shared token-side partial for each source position while separately projecting each computation's hidden state, avoiding repeated token-side projection work. On CNN and ShareGPT, our method achieves a 1.85–4.34 end-to-end training speedup over four existing training frameworks. Numerical checks show close agreement with SpecForge, and the trained drafters maintain comparable draft acceptance on held-out test examples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.