Positional Attention Signatures Distinguish AI-Generated from Natural Biological Sequences
Abstract
Biological foundation models generate increasingly realistic DNA and proteins, raising two questions: how do generated sequence distributions differ from natural ones, and how can those sequences be screened efficiently? We study positional attention variance (pAV), the variation of attention-weighted distance across positions in a frozen biological language model. Across six DNA generators, mean pAV is 53–84% of a matched natural reference under a multispecies Nucleotide Transformer. Preserving dinucleotide counts while shuffling order lowers median pAV by 16.2% across 200 natural windows, establishing arrangement sensitivity. A protein panel shows analogous differences under ESM2. In matched detection tests, inexpensive features are stronger on clean data, while standalone pAV retains macro AUROC 0.867 at 20% substitution, compared with 0.822 for compression. Adding pAV to a clean-trained compression classifier does not improve this stress-test performance. We separately evaluate DNA screening with multiscale sequence features and frozen observer representations, obtaining macro AUROC 0.995 in five-fold sequence-level cross-validation. Distilling the observer contribution enables a compact service with measured throughput of 334.6 sequences per second on one GPU. Together, these studies establish an arrangement-sensitive statistical probe and a computationally practical screening approach, with distinct measurement, evaluation, and calibration boundaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.