Robust Online Learning Against Biased LLM Judge with Budgeted Calibration
Abstract
Online learning is emerging as a powerful paradigm for adapting LLM systems, often relying on LLM-as-a-judge for feedback. However, such feedback can be systematically biased and may not faithfully reflect true task quality. Clean feedback can correct this mismatch, but is too costly to obtain at every round. This raises a fundamental challenge: how can an online learner leverage abundant but biased feedback with only a limited budget of clean feedback? We formulate this setting as a contextual bandit problem with an always-observed (biased) proxy reward and a limited budget for selectively acquiring clean feedback. We propose Paired-Audit Calibrated Exploration (PACE), which leverages paired proxy-clean observations to estimate true task quality and adaptively allocates the audit budget based on learned proxy-clean disagreement. We evaluate PACE on LLM routing, prompt selection, retrieval configuration selection, and synthetic benchmarks. Across our evaluations, PACE is most effective when proxy-clean disagreement is decision-relevant; on the four primary routing conditions, it reduces mean cumulative clean regret over proxy-only LinUCB by 36.5–58.8%. For theoretical characterization, we study PACE-U, a uniform-auditing variant that isolates paired proxy-clean calibration under linear realizability. PACE-U achieves clean regret, while any learner restricted to at most clean-feedback queries incurs regret on a two-dimensional instance, even with adaptive auditing. Thus, the dependence on the horizon and clean-feedback budget is minimax rate-optimal for fixed dimension.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.