Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
Abstract
Backdoored fine-tunes, jailbreaks and prompt injections all make a deployed LLM act against its deployer. We study one label-free signal for all three that needs no reference model, trigger knowledge or weight access. Layerwise Convergence Fingerprinting (LCF) scores how the last prefill token's hidden state moves between layers with per-layer diagonal Mahalanobis distances. It pools these with a shrinkage covariance and thresholds the result using 200 unlabelled in-domain examples, at under 0.1% overhead and before any token is generated. On four instruction-tuned models, one score separates backdoor triggers (per-architecture AUC 0.983–0.992) and jailbreaks whose template is held fixed and only the goal made harmful (0.83–0.97). It also flags all BIPIA text-payload injections. At a matched 5% false-positive rate (FPR), mean residual backdoor attack success is 2.8–8.9% (54 of 56 cells; 2 cells whose fine-tuned model already misbehaves on clean inputs are reported separately). End to end, LCF blocks 91.7% of jailbreaks that succeed without it and all 188 successful text-payload injections. The method also has clear limits. At 1% FPR, residual backdoor success is 10.6–32.6%; a detector-aware attacker evades it on 82.5% of Llama-3 GCG prompts; and a threshold calibrated on generic instruction data yields 70–100% FPR on email and code tasks. Our analysis traces these limits to two sources. The score mixes a harm component, which persists when surface form is held fixed, with a surface-novelty component that drives false positives. Anomalies also peak at threat- and architecture-specific depths that uniform pooling dilutes. LCF is therefore best used as a cheap first-stage screening signal for untrusted models, combined with further checks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.