acceptodds
Under review as a conference paper at ICLR 2027

Accelerating Decoding via Autoregressive-to-Parallel Distillation for Architecture-Agnostic Generative Recommendation

Abstract

Generative Recommendation (GR) retrieves items by autoregressively generating semantic identifiers, making beam-search inference costly. Existing speculative-decoding methods for GR either retrain the target with a draft head or train an external draft model while keeping the target frozen. The former incur substantial target retraining costs, whereas the latter still suffer from costly autoregressive draft inference and are limited by architecture-specific verification pipelines. To address these limitations, we propose the Autoregressive-to-Parallel Distillation framework for Generative Recommendation (APD-GR). This framework distills the frozen target's autoregressive prediction knowledge into a lightweight draft that computes SID position representations in parallel. Specifically, we design a two-stage distillation strategy that first establishes parallel prediction through SID supervision and then aligns the draft's prefix-conditioned distributions with those of the frozen target. To further support heterogeneous GR backbones, we propose a unified interface across decoder-only and encoder–decoder targets. In the experiments, to verify effectiveness and versatility of our proposed accelerating framework, we conduct extensive experiments across three datasets and five GR backbones. The results show that our method achieves an average end-to-end inference speedup of while largely preserving recommendation performance. To ease reproducibility, we release the code online https://anonymous.4open.science/r/anonymous-code-release-DC63.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.