When LLMs Answer with Ranges: Sample-Optimal Selection under Interval-Valued Responses
Abstract
Large language models (LLMs) are increasingly used as stochastic simulators to support selection from large sets of candidate designs and products. In many applications with high uncertainty, LLMs are more likely to return a numerical interval rather than a scalar. Motivated by the potentially vast number of candidates, we study large-scale selection under interval-valued responses. The objective is lexicographic: maximize the expected midpoint and, among alternatives that tie on midpoint, minimize the expected radius. We introduce nterval-esponse nterlaced election (IRIS) with two greedy algorithms. F-IRIS uses a ixed midpoint tolerance as an oracle benchmark; A-IRIS replaces it with an daptive candidate set. Theoretically, we prove that F-IRIS and A-IRIS are both sample optimal, that is, a budget linear in the number of alternatives achieves a positive probability of correct selection. By choosing larger finite initialization and budget coefficients, both methods can achieve any pre-specified accuracy. Experiments with Gaussian, uniform, and skewed midpoint noise support the asymptotic results. A study based on repeated Llama queries shows when interval information improves over midpoint-only selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.