acceptodds
Under review as a conference paper at ICLR 2027

SCORE: A Model Equality Test for Auditing Black-Box LLM APIs

Abstract

Users increasingly rely on black-box inference APIs as their primary interface to large language models (LLMs), for both closed- and open-weight models (e.g., Llama models are now widely accessed through Amazon Bedrock and Azure AI Studio), with little insight into which model is actually deployed. To reduce costs or maliciously alter model behavior, API providers may silently deploy quantized models or substitute entirely different ones, potentially degrading model performance and compromising safety. Detecting such substitutions is challenging because users cannot inspect model weights and, in most cases, lack access to internal model signals such as output logits, leaving generated text as the only observable. To address this problem, we propose Semantic Contrast for Output-only Rank-based Model Equality Testing (SCORE), an output-only auditing method that tests whether a target LLM API and its official reference API induce the same response distribution, thereby assessing whether the black-box API behaves consistently with the model it claims to serve. SCORE measures the change in semantic recovery induced by a control prefix, converts each target-model score into a randomized rank calibrated against reference-model scores, and tests the resulting ranks for uniformity at significance level \(\alpha\) using only generated text. We evaluate SCORE across multiple threat scenarios, including model quantization, harmful fine-tuning, and full model replacement, as well as on commercial inference API endpoints. Across these settings, SCORE consistently achieves higher statistical power than existing methods under stringent black-box constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.