acceptodds
Under review as a conference paper at ICLR 2027

PREFERENCE OPTIMIZATION IS NOT ENOUGH: HARNESSING CATALOG SUPPORT FOR GENERATIVE RECOMMENDATION

Abstract

Direct Preference Optimization (DPO) and its multi-negative extension, Softmax-DPO, align generative recommenders through relative item comparisons. However, unconstrained generation operates over the full autoregressive sequence space, where better preference margins do not guarantee that probability mass remains on valid catalog items. We term this failure mode catalog-support drift: a model can become better at preference discrimination while generating item strings that cannot be resolved against the catalog, undermining the practical usefulness of its recommendations. Addressing this mismatch requires learning not only which items users prefer, but also how to preserve support for valid outputs. To this end, we introduce , a dehallucination-oriented extension of Softmax-DPO that integrates sequence- and prefix-level support control with preference learning. At the sequence level, a full-slate ranking objective improves within-slate ordering, while a primal–dual constraint regulates the slate's total probability mass against an SFT-calibrated budget. At the prefix level, an on-policy trie loss penalizes probability assigned to catalog-invalid continuations along the evolving policy's generation paths. Together, these objectives shape the policy itself without imposing a catalog mask at inference. Under full-policy ancestral sampling, our analysis relates slate probability mass to exact outside-slate risk and bounds finite-horizon catalog-exit probability by the expected prefix-support loss. Experiments across five real-world recommendation datasets reveal the practical severity of catalog-support drift. Extensive results show that DeHall achieves superior performance on both catalog validity and policy-ranking accuracy, reducing nucleus-sampling catalog-string violation rates. Component analyses further support the complementary roles of prefix- and sequence-level support control. Code is available at https://anonymous.4open.science/r/DEHALL-SDPO-6776/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.