acceptodds
Under review as a conference paper at ICLR 2027

The Cost of Comparing Coupled Generators

Abstract

Sharing randomness between two generators can make their outputs agree, reducing the work needed to compare them. Directly judging the coupled outputs, however, generally changes the win probability being estimated. We study how agreement can reduce evaluation cost while retaining the independent-output win probability. Let be the average probability that the coupled outputs differ. We give a procedure that does not know and, at accuracy and confidence , uses expected generation and judgment costs of order and , respectively. When , both dependences are necessary on the same worst-case instance, even for comparators induced by a strict scalar ranking and algorithms that adaptively judge outputs they have not generated. A pilot based only on output identity provides the unknown mismatch scale. The analysis connects crossed paired comparisons and reference differences, two ways to retain the target and exploit agreement. Recorded-pool experiments with language models on two text-generation tasks show 29–49% lower RMSE for crossed comparisons than independent pairing at the same generation count, with fewer judgments on all four pools.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.