acceptodds
Under review as a conference paper at ICLR 2027

NCCA: Nested Credit with Conditional Amplification for LLM-Based Conversational Recommendation

Abstract

Conversational recommender systems often use past dialogues to recover user preferences missing from a current request. We study conversational recommendation as a two-stage pipeline using large language models (LLMs) that first distills the dialogue history into a request-specific user state in natural language and then reranks item candidates from that state. Because downstream ranking quality can be computed automatically from observed feedback, the pipeline can be trained end-to-end with reinforcement learning with verifiable rewards (RLVR). A straightforward implementation assigns the same outcome advantage to both stages. In conversational recommendation, however, this signal mixes two sources of variation: state quality and stochastic ranking execution, with the latter especially strong when only a few candidate positions yield high reward. This coupling makes the feedback for each stage noisy: state updates inherit variation from stochastic ranking execution, while reranking updates vary with state quality. To solve this problem, we propose NCCA (Nested Credit with Conditional Amplification), a credit-assignment strategy for joint two-stage training. In NCCA, Nested Credit samples multiple states per request and multiple rankings per state, averages sibling rewards for state credit, and performs within-state relative credit for reranking. Thus, it suppresses execution noise upstream and isolates state-conditioned ranking variation downstream. Furthermore, Conditional Amplification upweights positive within-state advantages while retaining negative corrective updates. This strengthens the learning signal from the relatively few successful rankings without discarding corrections from lower-reward rankings. Our experiments show that NCCA improves NDCG@10 by 13.2% and 15.1% over the shared-advantage RLVR baseline on TG-ReDial and LLM-REDIAL, and enables a 4B model to match strong general-purpose LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.