RMF: Runtime Model Fingerprinting for Detecting API Model Substitution
Abstract
Third-party LLM aggregation platforms provide convenient access to multiple models through unified APIs, yet their opaque backend makes it difficult for users to verify which model actually serves a request, enabling API fraud through model substitution. Existing black-box fingerprinting methods primarily perform pre-deployment verification and therefore cannot reliably authenticate the model during actual task execution. To address this limitation, we propose a runtime model fingerprinting method (RMF) that embeds lightweight fingerprint probes directly into real-world task prompts. RMF constructs compact behavioral fingerprints from model-specific preferences over underdetermined ranking tasks and employs an adaptive probe search framework to select discriminative fingerprints for different task contexts. We evaluate RMF on 23 frontier LLMs across four real-world application domains. Compared with state-of-the-art baselines, RMF substantially improves model identification in both binary verification and multiclass identification settings, while introducing no statistically significant degradation in task performance and negligible additional token overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.