acceptodds
Under review as a conference paper at ICLR 2027

The Statistical Benefits of Multiple Responses for Learning from Demonstrations

Abstract

Many generative systems return multiple candidate responses and are evaluated according to the best one. Recent work shows that, when demonstrations are optimal, pass@ can reduce the sample complexity of learning from demonstrations by a logarithmic factor in . We ask what happens when the demonstrator is not assumed to be optimal. We find that multiple responses provide a qualitatively stronger benefit in this setting. In a finite reward-class model with no reward feedback, moving from pass@ to any pass@ with changes the worst-case dependence on target accuracy from to , uniformly over demonstrator quality. Under standard evaluation, where an unknown reward is fixed before training, increasing provides an additional and distinct benefit: the optimal dependence on a reward class of size improves from to . We further show that these two effects can be separated. Under robust evaluation, where one learned policy must compete with the demonstrator simultaneously for every reward in the class, the fast dependence persists, while the improvement can disappear. We establish matching upper and lower bounds in the corresponding regimes and give a greedy multiplicative-weights learner achieving the upper bounds without any assumption on demonstrator quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.