MERA: Matrix-Entropy Reranking Actor-Critic for Sequential Recommendation
Abstract
Sequential reranking is a critical final stage in recommender systems. Given candidates and display positions, it makes sequential decisions over discrete ordered slates. Existing methods often optimize proxy scores, fusion weights, or latent controls before decoding them into slates, creating an action-interface mismatch. We propose Matrix-Entropy Reranking Actor-Critic (MERA), which models reranking as a Markov decision process with display-or-discard assignments as actions. Using the item–slot assignment probabilities of ranking policies, we define matrix-entropy-regularized Bellman operators and establish their contraction properties. We characterize the marginal actor problem's first-order stationarity through a KKT decomposition into an inner entropic transport problem and an outer projected score-alignment condition. MERA parameterizes feasible assignment plans using scores and solves the inner problem with Sinkhorn scaling. Critic action gradients at hard assignments sampled by Gumbel perturbation and Hungarian rounding guide projected score alignment. On two datasets with three simulators, MERA improves mean total session reward by up to 14.43% over the strongest of seven baselines per setting, with competitive session-level exposure diversity. MERA maintains low inference latency as candidate sets grow, supporting deployment in industrial reranking systems. Our code is available at https://anonymous.4open.science/r/MERA_code-CD4D/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.