LayerTrace: Learning Source-Specific Token Latents for Sparse Attention
Abstract
Long-context decoding repeatedly reads historical keys and values, while sparse attention adds work to identify useful entries. LayerTrace moves that decision to an earlier transformer layer and learns the token representation used to make it. Each source stores a latent for every historical token, obtained from its source-block output through trainable content and positional projections. Independent destination-layer and query-head scorers use the current source output to prepare future read maps. Encoders and scorers are trained together against dense attention distributions and outputs, with a frozen backbone; destinations attend to selected original keys and values. On 503 LongBench v2 questions, Qwen3-4B at threshold 2 answers 157 correctly versus Dense's 153 while selecting 31.58% of original KV. In separate fixed-input timing at the same operating point, 32K context and batch 16, whole-decoder time falls from 62.68 to 36.34 ms per step. A 56-question DeepSeek development diagnostic scores 59.08 versus 49.40, but all-selected numerical discrepancies remain unresolved. Higher thresholds can collapse quality, and single-request execution can remain slower. These exploratory observations do not yet establish quality-preserving speedups under matched batched answer generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.