acceptodds
Under review as a conference paper at ICLR 2027

Stemma: Induced Decision Regions Reveal LLM Provenance

Abstract

LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely rely on evidence tied to particular elicitation interfaces or response-level characteristics, but these signals may shift under prompting, adaptation, or deployment without changes in model lineage, weakening provenance reliability. To address this limitation, we introduce induced decision regions by mapping open-ended outputs into a shared finite decision space, abstracting away surface-form variation and reframing provenance testing around induced decision region inheritance. Empirical analysis shows stronger inheritance in related models and reveals that the provenance signal is unevenly distributed across probes, with discriminative power concentrated in a small subset whose informativeness transfers across model pairs. Building on this formulation, we propose Stemma, a black-box LLM fingerprinting method that estimates induced decision region inheritance using probes selected for stability, robustness, and specificity. Across 770 source-suspect pairs from 56 public checkpoints spanning diverse model transformations, Stemma achieves 0.967 AUC and 87.1% TPR at 1% FPR. It further achieves 0.995 AUC and 92.4% TPR at 1% FPR on 1,260 pairs covering 91 deployment instances, demonstrating robustness to deployment variations. Across both benchmarks, Stemma substantially outperforms four representative baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.