acceptodds
Under review as a conference paper at ICLR 2027

From Global Budgets to Per-User LTV Health: Offline Reinforcement Learning for Personalized Coupon Allocation under Extreme Population Skew

Abstract

Coupon incentives are widely used to promote purchases and sustain user engagement. However, existing systems typically apply uniform rules within user tiers or optimize aggregate returns under a global budget, overlooking differences in users' lifetime value (LTV). This limitation is particularly severe in e-commerce referral platforms dominated by low-activity users, where sparse histories and highly skewed populations further complicate LTV estimation and offline policy learning. To address these challenges, we formulate daily joint coupon allocation as an offline reinforcement learning problem for personalized fee-rate control with user-specific, LTV-aware penalties. We propose a group-relative offline RL algorithm built on implicit Q-learning, termed LIGER, that represents evolving LTV health through time-varying state features and a structured reward balancing immediate incentive efficiency with long-term value changes. To mitigate majority-group dominance, LIGER combines group-balanced sampling with within-group reward and advantage normalization. It further replaces a separately fitted state-value network with a temperature-controlled softmax Bellman target derived from twin Q-critics, avoiding an additional value-function fit on imbalanced data. Offline evaluations and online A/B tests in a real-world e-commerce setting demonstrate the effectiveness of LIGER.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.