acceptodds
Under review as a conference paper at ICLR 2027

To Defer or Not to Defer? Online Learning with LLMs and Congested Experts

Abstract

Modern AI systems increasingly combine large language models (LLMs) with human experts, using the LLM to handle routine queries while deferring uncertain cases for expert review. When expert capacity is limited, however, deferral decisions must account not only for the reliability of the LLM but also for congestion in the expert queue. We formulate this as an online sequential decision-making problem in which, for each query, the system observes the confidence signal of the LLM's answer and must decide whether to output the answer or defer the query to a capacity-constrained expert. The reliability of the confidence signal is initially unknown and must be learned from expert feedback, which is both selectively observed (only deferred queries are labeled) and delayed by queuing. We propose KL-PBD, a learning algorithm that uses Kullback–Leibler lower confidence bounds on the unknown reliability parameters to guide deferral decisions. For confidence levels and horizon , KL-PBD achieves gap-dependent regret when the accuracy at every confidence level is bounded away from the deferral boundary, and worst-case regret. We establish matching lower bounds, showing that these rates are optimal up to logarithmic factors. Finally, we validate our approach on synthetic instances and real-world question-answering datasets, namely GSM8K, NuminaMath, and MedMCQA, using Qwen LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.