AuditGym: Can Black-Box Audits Verify the Model Behind an LLM API?
Abstract
Large language models (LLMs) are increasingly accessed through APIs and embedded in agent systems, yet users have limited evidence that the API actually serves the model it claims. Existing black-box behavioral audits approach this problem through equality testing, model identification, and fingerprinting. However, their outcomes are difficult to interpret and compare: a detected discrepancy may reflect either a different model or a change in how the same model is deployed. We introduce AuditGym, a benchmark for systematically evaluating LLM identity verification under controlled deployment changes. AuditGym separates model specification, including checkpoints, learned adapters, and quantization, from generation policy and the interaction layer, and tests whether auditors reject changes to the former while accepting changes to the latter. We evaluate four auditing approaches across seven change families and 19 checkpoints, including models running inside real agent harnesses. All evaluated methods achieve only limited reliability for the identity verification: all four detect checkpoint substitution but miss finer model changes such as adapters and quantization. More seriously, every method rejects unchanged models under certain generation, watermarking, system-prompt changes, or agent harnesses. With these results, we reveal that current auditors confuse model changes with deployment changes, and AuditGym provides a controlled foundation towards closing this gap. Our anonymous codebase is available at https://anonymous.4open.science/r/auditgym_iclr-BC1E/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.