Regulating Human-in-the-Loop Learning for Dynamic LLMs
Abstract
Large language model (LLM) platforms increasingly host multiple models whose response quality evolves through user interactions and online adaptation. This creates a human-in-the-loop model-selection problem in which myopic users favor the best immediate option, while the platform seeks to maximize long-run user satisfaction. We formulate this problem as a parametric stochastic rising bandit with myopic users, using %low-dimensional saturation families to capture adaptation-driven reward growth. For a platform-controlled benchmark, we propose GAP-UCB, a growth-aware prospective UCB algorithm that fits reward trajectories via nonlinear least squares and compares models using optimistic prospective rewards. %Under the parametric formulation, GAP-UCB achieves \(O(T\log T)\) regret, improving over the existing \(O(T^2/3)\) guarantee for general stochastic rising rested bandits. We then return to the user-driven setting and develop iGAP-UCB, an ex-post budget-balanced payment mechanism that regulates users toward the platform-preferred action while preserving the same regret order up to a fee-dependent factor. Experiments on real LLM fine-tuning trajectories validate the effectiveness of GAP-UCB in the platform-controlled setting and show that iGAP-UCB closely tracks the platform-preferred benchmark under budget-balanced user regulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.