acceptodds
Under review as a conference paper at ICLR 2027

Same Artifact, Two Faces: Why Black-Box Lineage Fingerprints Look Both Robust and Fragile

Abstract

Knowing which training branch a deployed model is produced from matters for model auditing and policy. We ask to what extent the post-training branch of same-family models can be identified from black-box fingerprints alone, i.e. output texts. Qwen3 base/IT pairs are the most comparable case: under a pre-training continuation format, both branches generate readable text. We collect 100 to 200 MATH500 problems per model and train a 16-feature output classifier. Within a single scale the separation is near-perfect (0.941 / 0.947), but only when the reading is uncontrolled: it is not stable across the five scales we ran, and it moves further than the headline effect we report when we change the prefix length we keep, chosen arbitrarily. Once length and problem difficulty are controlled together, the leave-one-model-out transfer is no longer distinguishable from chance, and it does not collapse to a clean negative either. Disguise, prompt and answer rewriting, template and temperature perturbation, and cross-lingual attacks all leave the reading without a discernible decrease, because the signal carrying it is termination and length; the one arm that moved the reading moved it by shortening the output, which is that channel.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.