acceptodds
Under review as a conference paper at ICLR 2027

Same Target, Different Paths: Learning from Multiple Reasoning Trajectories for Generative Recommendation

Abstract

Reasoning-based generative recommenders can support both explicit reasoning and direct Semantic ID (SID) generation within the same model. We study how multi-trajectory reasoning post-training affects this hybrid inference setting. In self-sampled rejection-sampling fine-tuning (RFT), multiple successful trajectories for the same request provide diverse reasoning paths while terminating at a shared target SID, coupling richer reasoning supervision with repeated terminal-target supervision. We introduce target-masked RFT (TM-RFT), which removes the terminal-target loss from augmentation trajectories while preserving supervision on their reasoning traces and on ordinary SFT data. On the OneReason benchmark, TM-RFT with up to four trajectories per request achieves a Weighted Pass@64 of 0.7576, compared with 0.6912 for SFT-only and 0.7011 for a group-normalized control that rescales the repeated target loss. Answer-supervised variants achieve higher standalone reasoning and direct-generation scores, but exhibit substantially stronger agreement between the two inference modes, with up to 44% of their candidates shared. In contrast, TM-RFT maintains route separation close to SFT-only while retaining 99.6% of the strongest single-mode performance under a fixed 64-candidate budget. On three public Amazon datasets, TM-RFT further improves two-route hit rate over both the original checkpoints and SFT-only.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.