From Detecting to Characterising LLM API Updates: Black-Box Inference of the Change Mechanism
Abstract
Recent works proposed the task of change detection of LLMs behind APIs. This responded to practical concerns that some providers were observed to modify their models (e.g., with fine tuning or system prompt) without warning their end-users, and that caused potentially severe regression on downstream applications. This paper proposes to go a step further than mere change detection: we study change-type inference. Given an LLM endpoint that is known to have changed, we predict whether the update involved fine tuning, a system-prompt modification, or abliteration, using only sampled responses from the updated endpoint. We first give a theoretical grounding for the possibility of such a novel task, by demonstrating that changes that are local to a model provoke a monotonically evolving distance from that base model, which is traceable in its output by black-box auditing methods. We then propose (Change-Type Inference), a scheme that decides which change has been performed, and that relies on Border Inputs, which are sensitive to model modifications and already leveraged in change detection problems. All in all, this paper demonstrates the feasibility of detecting which change occurred on an LLM in a black-box setup, which we believe will be of interest for auditors and regulators.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.