Representation-Guided Preference Learning for Multi-Episode Group Decision-Making
Abstract
We study representation-guided group decision-making with a fixed pool of participants over a sequence of episodes. During each episode, a group coordinator observes scalar utility feedback from each participant in response to text descriptions of options, and seeks to find a fair solution to the posed decision problem. We formalize the setting as a bilinear low-rank bandit and apply an approach of alternating minimization with LCB-minimax exploration. We prove a cumulative pseudoregret bound that scales with the spectral dimension of the representation , rather than the ambient encoder dimension . On a controlled multi-episode simulator, the algorithm reduces the total episode-best-plan gap to vs –; per-episode error also drops sharply: from in episode 1 to in episode 3, and thereafter remains roughly in the range. The same reductions carry over to real pretrained encoders (Qwen3-8B, Llama-3.1-8B-Instruct, gemma-3-4b-it, and all-MiniLM-L6-v2), where jointly learning the shared probe and per-user preferences reduces by 1.9–2.5 versus random projections and 3.2–4.9 versus one-hot features.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.