acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Sensitivity Probes: Distilling Popularity-Robust Judgment into Fast Sequential Recommenders via GRPO

Abstract

Sequential recommenders trained on implicit feedback inherit popularity bias, and while LLMs can reason explicitly about this confound, such reasoning is too slow for production serving — a claim we verify by direct measurement rather than assumption. We introduce counterfactual sensitivity probes — real item pairs matched on content but deliberately mismatched on popularity — as both a training signal and a diagnostic for whether a model's preference tracks content or popularity, and use them to distill a large teacher's popularity-debiasing judgment into a small, non-reasoning student via two GRPO recipes: single-stage training on the full teacher-labeled mix, and a two-stage curriculum that warms up on the diagnostic pairs first. Along the way, we fix a previously undiagnosed unbounded policy-entropy growth in GRPO training via an adaptive KL-to-reference controller and a capped IPS reward, enabling one pre-registered stopping rule applied uniformly to both recipes and two classical baselines (alignment-only, IPS-weighted), and we pair the resulting accuracy benchmark with a deployment-oriented cost analysis, measured rather than assumed: the distilled student is faster and higher-throughput than the teacher, while curriculum's two-stage training costs a single-stage run, driven by job-launch overhead rather than extra computation. On Amazon Beauty, results are honestly mixed: curriculum is the only method that clears both baselines (68.79% sensitivity accuracy).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.