acceptodds
Under review as a conference paper at ICLR 2027

Will LLMs Eventually Become Indistinguishable? Rethinking Authorship Attribution

Abstract

LLM authorship attribution aims to identify the source model of a generated output by leveraging specific characteristics in its outputs. Existing approaches implicitly assume that different LLMs present sufficiently distinguishable generation patterns. However, this assumption may be challenged by knowledge distillation, which explicitly transfers knowledge between models while potentially altering the source-specific signals used for attribution. In this work, we systematically study how knowledge distillation affects black-box LLM authorship attribution. Across two benchmarks, nine source classes, and seven representative attribution methods, we find that knowledge distillation consistently impairs authorship attribution, causing distilled student outputs to become increasingly confused with their corresponding teachers. Moreover, we find that retraining attribution models on post-distillation outputs recovers only part of the lost attribution accuracy, indicating that the degradation cannot be fully explained by distribution shift. Instead, knowledge distillation appears to reduce the intrinsic distinguishability between student and teacher outputs. To theoretically explain this phenomenon, we develop an information-theoretic framework showing that knowledge distillation reduces the mutual information between generated text and its source, and further characterizing how much of this source information remains accessible to an attribution model. Together, our findings reveal an overlooked challenge for LLM provenance in an increasingly interconnected LLM ecosystem.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.