acceptodds
Under review as a conference paper at ICLR 2027

Explainability-Guided History Distillation for Efficient Sequential Recommendation

Abstract

Sequential recommendation models, spanning self-attention based Transformer architectures to state-space models, rely on processing complete interaction histories, incurring inference costs that scale with sequence length. While efficient attention variants and heuristic truncation mitigate this computational footprint, they overlook internal model attribution signals and neglect the efficiency gains achievable by distilling the raw interaction sequence itself. To bridge this gap, we propose explainability-driven sequence distillation, a retraining-free framework that leverages post-hoc attribution methods to score each interaction's predictive contribution. By retaining only the highest-scoring positions, our framework compresses user histories offline, enabling shorter sequences to accelerate online inference without altering underlying model parameters. Across four dataset configurations and four architectural paradigms, including autoregressive, bidirectional, state-space, and long-sequence backbones, Integrated Gradient-based sequence distillation preserves or improves ranking quality while substantially reducing inference GFLOPs, outperforming recency-based truncation by an average of 11.5% across all evaluated settings. Crucially, even on benchmarks with extreme temporal recency bias, our approach matches or exceeds recency truncation in the performance-efficiency trade-off, an advantage that widens under noisy input sequences. These findings demonstrate that post-hoc explainability, traditionally restricted to model interpretation, serves as a general, retraining-free principle for distilling raw input contexts and reducing online serving costs in sequential recommendation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.