Optimistic Preference Exploration with Finite Reward Ensembles
Abstract
Learning systems increasingly choose the feedback and experience used to improve them. In preference learning, a comparison can favor a strong current response or reveal information that helps future decisions. We develop Optimistic Learning with Ensembles (OLE), a framework connecting finite reward sampling, comparison selection and feedback updates. Its central idea is to assess reward ensembles through the learning decisions they induce. Our analysis establishes a query-regret guarantee for optimistic posterior sampling in a bounded linear preference model. An exact construction explains when information acquired by exploration repays its immediate cost and how dependence between reward samples changes this tradeoff. For implementation, we identify the relative score errors that govern the precision of a chosen selection rule. Controlled learning studies examine the interaction between sampling and selection, while neural experiments demonstrate more precise implementation at a fixed reward-evaluation budget. Agentic OLE applies reward-sensitive selection to generated training trajectories, with evidence from question answering, research and interactive tasks. Together, the results connect preference exploration to the broader design of learning systems that choose their own experience.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.