acceptodds
Under review as a conference paper at ICLR 2027

QUIET: QA-Free Unbiased Implicit Preference Extraction with Semantic F1 Rewards

Abstract

Personal LLM agents rely on preferences extracted from user–agent interactions to provide personalized responses. However, such preferences are often expressed only implicitly across lengthy interactions, making reliable extraction challenging. Although specialized small-scale extractors offer an efficient solution, we observed that question-answering (QA) rewards used in reinforcement learning (RL) can bias extractors toward QA-specific patterns at the expense of preference fidelity. Specifically, even on the rigorous PersonaMem-V2 benchmark, the QA-reward-trained model performs worse when conditioned on its own extracted preferences than without them, highlighting the need for improved extraction methods and new benchmarks that do not overlook such biases. To this end, we introduce FairPersonaMem-V2, a bias-aware reconstruction of the benchmark, and QUIET (QA-free Unbiased Implicit preference ExTraction), a reward design for RL to train a preference extractor directly from ground-truth preferences, without requiring QA supervision. On FairPersonaMem-V2, the small-scale Qwen3-4B-Instruct extractor trained with QUIET achieves a 32% relative MCQ accuracy gain and an 11 higher semantic F1, with 21 fewer training context tokens compared to the QA-reward trained baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.