READER: Dynamic LLM Provenance from Query-Varying Interactions
Abstract
Auditing the provenance of large language models (LLMs) is essential for verifying service claims and detecting model substitution. Yet API services rarely expose the internals required by white-box methods, while many black-box methods establish comparability through shared diagnostic prompts. Auditing query-varying interactions therefore requires comparable source evidence as observations accumulate. We formalize this setting as dynamic black-box LLM provenance within a fixed set of enrolled models. To make these interactions comparable, we introduce READER, which uses a frozen proxy LLM as a common measurement space and encodes the mean and coarse evolution of activations along each response into spectral fingerprints. A linear probe trained once at enrollment extracts source evidence that Bayesian accumulation combines across observations. On Agent500, our benchmark containing 50,000 responses from 100 models to 500 heterogeneous agent prompts, READER reaches attribution accuracy from one response and from 100 on unseen prompts, exceeding fine-tuned DeBERTa and LLM-DNA with frozen sentence encoders at the respective budgets. Experiments across proxy families further demonstrate strong cumulative attribution, while static analyses reveal model relationships in the same spectral representation. READER thus turns ordinary interactions into cumulative evidence for provenance auditing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.