PREMO: Complementary Training For Personalized Preference Discovery
Abstract
Personalized decisions require discovering query-relevant user preferences from long, noisy histories. These preferences can be summarized in a memo that covers relevant preferences while remaining faithful to historical evidence. However, the complete set of relevant preferences is unobserved, making memo discovery difficult to supervise. We introduce PREMO (PREference MemO), which alternates a single model between two complementary roles. As a proposer, it discovers new atomic preferences and is rewarded by independent verification against the history. As a memo policy, it generates a memo from the query and history alone, rewarded for covering verified preferences without asserting rejected ones. Proposer training expands the verified supervision and directly trains preference recognition, while memo training learns to express these preferences in a useful summary. Our theoretical analysis identifies conditions under which proposer training followed by memo training improves both coverage and faithfulness over memo-only training. We evaluate PREMO on shopping, movie recommendation, and everyday decisions. A trained 4B model writes memos of about 170 tokens that effectively summarize histories 40 to 100 times longer. It improves coverage and faithfulness over memo-only training and raises downstream accuracy by up to 3.7 points. Its memos also match or beat its frontier-model teacher’s own: PREMO covers more and raises downstream accuracy by up to 6.8 points, even when GPT-5.4 itself decides. The advantage over memo-only training holds at 35B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.