acceptodds
Under review as a conference paper at ICLR 2027

Reliable LLM Provenance via Panic-Pattern Fingerprinting

Abstract

Training competitive large language models (LLMs) is expensive, making LLMs high-value intellectual property and motivating reliable provenance tests for both open-weight and API-served models. In this paper, we propose a unified provenance framework based on *panic patterns*—model-specific fingerprints elicited by stressful inputs. In the white-box setting, we introduce **P**anic **A**lignment **D**ifference (PAD), which fingerprints models by measuring stress-induced shifts in layerwise representation alignment via a clean-versus-adversarial differential. In the black-box setting, we introduce **O**pen-ended **Q**uestion **T**est (OQT), which uses open-ended questions as stressors and compares residual response embeddings after removing the consensus answer estimated from a reference pool. Across a range of open and proprietary LLMs, our stress-response fingerprints reduce spurious similarity, remain informative under the evaluated model variations, and provide complementary provenance evidence across access levels with bounded verification cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.