One Character at a Time: Frustratingly Simple and Cost-Efficient Black-Box Model Equality Testing
Abstract
Users of third-party large language model APIs may have little direct evidence that the served model matches the advertised one. Black-box model equality tests offer a practical auditing mechanism, but existing methods can require many queries, long generations, or a costly query optimization stage. We study whether reliable auditing is possible with minimal resources and find a frustratingly simple yet effective equality-testing strategy that requests only one character per model response. Our method asks both the candidate and a trusted reference model to output one symbol from a fixed alphabet. Repeated responses yield empirical categorical distributions, which we compare using total variation (TV) distance and a paired permutation test. We evaluate our method using a cost-versus-statistical-power protocol, with same-model and different-model comparisons to assess Type I and Type II errors. Comprehensive empirical evaluation shows that our method achieves strong performance at substantially lower cost than existing state-of-the-art methods. We also develop a theoretical analysis of when single-symbol responses suffice for model equality testing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.