acceptodds
Under review as a conference paper at ICLR 2027

ZeroGR: Constraint-Driven Execution for Generative Recommendation

Abstract

Constrained decoding is essential for generative recommendation, ensuring that generated semantic identifiers correspond to valid catalog items. At large catalog scales, existing approaches incur substantial overhead in constraint lookup and index storage. Even when constraint lookup is accelerated, applying constraints only after output projection wastes computation on tokens that cannot lead to valid recommendations. We present \sys, a constraint-driven inference system that uses catalog information before projection to restrict scoring to valid continuations. \sys represents catalog constraints as heterogeneous per-step mappings that directly expose valid next tokens and their successor states, avoiding repeated prefix traversal and explicit successor storage. A parallel streaming construction method builds these mappings from sorted semantic identifiers while overlapping CPU processing with transfer to the GPU. Specialized GPU kernels schedule valid projections through a compact work queue and reuse hidden vectors across candidates, avoiding both invalid projections and full candidate compaction. On a real-world recommendation dataset, \sys reduces projection-and-selection latency by 34.8% compared with representative constrained-decoding approaches. Deployed on a production recommendation platform serving 400 million daily active users, \sys reduces index memory usage by 66.2% compared with STATIC and end-to-end latency by up to 61.0% compared with the previously deployed CPU Trie serving approach.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.