acceptodds
Under review as a conference paper at ICLR 2027

MERA: Matrix-Entropy Reranking Actor-Critic for Sequential Recommendation

Abstract

Sequential reranking is a critical final stage in recommender systems. Given candidates and display positions, it makes sequential decisions over discrete ordered slates. Existing methods often optimize proxy scores, fusion weights, or latent controls before decoding them into slates, creating an action-interface mismatch. We propose Matrix-Entropy Reranking Actor-Critic (MERA), which models reranking as a Markov decision process with display-or-discard assignments as actions. Using the item–slot assignment probabilities of ranking policies, we define matrix-entropy-regularized Bellman operators and establish their contraction properties. We characterize the marginal actor problem's first-order stationarity through a KKT decomposition into an inner entropic transport problem and an outer projected score-alignment condition. MERA parameterizes feasible assignment plans using scores and solves the inner problem with Sinkhorn scaling. Critic action gradients at hard assignments sampled by Gumbel perturbation and Hungarian rounding guide projected score alignment. On two datasets with three simulators, MERA improves mean total session reward by up to 14.43% over the strongest of seven baselines per setting, with competitive session-level exposure diversity. MERA maintains low inference latency as candidate sets grow, supporting deployment in industrial reranking systems. Our code is available at https://anonymous.4open.science/r/MERA_code-CD4D/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.