PROVENANCELAB: AN AUDITABLE LONGITUDINAL CORPUS FOR LLM WATERMARK TRANSITION ANALYSIS
Abstract
Major providers began embedding machine-readable watermarks in large language model (LLM) text output during 2026 amid the implementation of the EU AI Act's synthetic-content marking requirements, but for many pre-existing commercial models the exact date at which outputs on a particular serving surface begin carrying a watermark is not publicly disclosed, and marking is being rolled out gradually. This creates a narrow and irreproducible window: outputs a model produced before it was marked cannot be regenerated afterward. We present ProvenanceLab, a longitudinal, externally auditable corpus of Claude, GPT-5.6, and Gemini outputs collected across this deployment boundary. Each record carries cryptographic integrity hashing and, for the strongest cohort, service-generated cloud evidence (invocation logs, provider request identifiers), retention-protected storage, and external RFC 3161 timestamps, so that the record of what a model emitted during the transition survives in a form a third party can audit. We further validate the statistical instrument that a future analysis will use: a controlled watermark on/off calibration experiment with the SynthID-Text logits processor on Gemma-2B-IT (1,460 generations) shows, under prompt-clustered inference and prompt-disjoint held-out evaluation, that pre-specified lexical and bigram diversity features register a known watermark with reproducible on/off effects (Cliff's ; 95% cluster-bootstrap CI excluding zero) that strengthen in longer outputs ( for 250–500-word outputs) and persist on held-out prompts. The corpus, the provenance architecture, and the calibrated instrument are the contribution and are complete in this paper; a provider-supported verifier, when one becomes available, supplies the final retrospective label. We do not claim watermark detection or attribution for any commercial model from text statistics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.