TraceAttention: Efficient Long-Context Attribution in Reasoning LLMs via Cache-Preserving Evaluation
Abstract
Context attribution explains how input spans influence model responses using internal-state analysis or perturbations. Perturbation-based methods score a fixed response under block masks of the same long context, but physical deletion requires repeated, expensive suffix reconstruction even with prefix caching. We introduce TraceAttention, a training-free evaluator that prefills once and reuses the full-context state by masking only target-to-context attention. This preserves the mask-to-score interface of existing attribution estimators while replacing repeated context reconstruction with target-only replay. The intervention differs from deletion, so we evaluate its utility using held-out physical deletions. For reasoning LLMs, scoring the full reasoning and answer exposes the reasoning's context dependence but increases per-mask work. Dual TraceAttention (DualTA) addresses this cost by selecting answer-relevant reasoning segments for a compact target. Across 15 model-context settings, TraceAttention and DualTA provide geometric-mean incremental speedups of 9.6 and 14.2 over prefix-cached deletion, respectively, while retaining over 90% of the deletion reference on average for both the linear datamodeling score (LDS) and Top-5 removal utility. On Qwen3.5-27B at a mean context length of 205K tokens, DualTA reduces incremental attribution latency from 841.0 to 13.1 s, a 64.2 speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.