acceptodds
Under review as a conference paper at ICLR 2027

Provable LLM Tampering Detection: Catching Model Provider Tampering of Open-Weights LLMs

Abstract

Prior theory-driven methods for detecting black-box LLM configuration tampering do not preserve their guarantees of a low false-positive rate (FPR) when the values of multiple common sampling parameters—such as temperature, token logit bias, and top-p—are hidden. Such sampling parameters can unpredictably and nonlinearly distort LLM output distributions without changing the underlying model configuration, making benign variation in sampling parameters difficult to statistically distinguish from genuine model configuration mismatches such as those arising from model substitution, secret fine-tuning, prompt injection, and weight quantization. To our knowledge, our paper establishes the first tamper-detection method for black-box LLM configurations that formally guarantees a low long-run FPR in the presence of a broad set of hidden standard sampling parameters. To this end, we develop a general theoretical framework for tamper-detection in this setting, and then construct two concrete classes of tests—the Temperature Test and the Logit Bias Delta Test—each provably achieving a long-run FPR of 0%. In our follow-up demo spanning four major tampering classes—model substitution, secret fine-tuning, system prompt injection, and weight quantization—the Temperature Test achieves 100% true-positive rate (TPR) and 0% FPR, while the Logit Bias Delta Test achieves 92.9% TPR and 0% FPR. Estimating the cost of each test in the demo at the prevailing rates of currently popular open-weight model APIs places the upper bound between 2.6, with our adaptive early-stopping technique further reducing the average estimated cost of a positive test to 0.6.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.