Bridging Exploration and Persona: Information-Directed Dueling for Test-Time Agent Personalization
Abstract
Large language model (LLM) agents have achieved remarkable success, yet they often lack the flexibility to adapt to individual user preferences. While recent frameworks like PersonaAgent attempt test-time alignment, they face two critical bottlenecks: (1) Heavy Data Dependency: they rely on pre-annotated historical preference data (ground-truth responses) for optimization, which is often unavailable in realworld scenarios; and (2) Efficiency Constraints: the highdimensional search space of personas leads to “interaction fatigue”, where users are burdened by excessive feedback requests. In this paper, we propose a Personalized Neural Dueling Bandit framework that enables seamless test-time personalization using pairwise preference feedback without preannotated target responses or scalar utility labels. We introduce an Information-Directed Exploration strategy inspired by Information-Directed Sampling (IDS) to navigate the persona space efficiently. By optimizing an information ratio that balances the estimated utility gap against the expected information gain, our method strategically selects the most informative comparison pairs to query for user preference feedback. Experimental results demonstrate that our approach consistently outperforms baselines in both personalization accuracy and interaction efficiency, effectively
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.