Black-Box Forensics for Conversational LLM Agents
Abstract
LLM-powered scams are proliferating, API providers silently change the models behind their endpoints, and jailbreaks are increasingly model-specific. Each of these calls for black-box forensics of conversational LLM agents: learning what powers an agent from nothing but ordinary, non-adversarial conversation with it. We study two capabilities. Attribution identifies which of a known set of base models or system prompts powers an agent, and we treat it as classification: a classifier trained on stylistic and n-gram features identifies the base model with 98% accuracy from a few turns of conversation. System prompts can be attributed the same way, but only by collecting labeled conversations for every candidate prompt, and the prompts deployed in the wild are too many and change too often for that. Fingerprinting instead decides whether two conversations come from the same agent, and so needs no training data from the agents it is applied to. Our cross-encoder scores a pair of conversations zero-shot: on system prompts never seen in training, it reaches an AUC of 0.768 from a single pair of conversations, and 0.943 when 50 conversations per agent are aggregated. It remains useful when every reply is paraphrased by another model, and it gives positive results on real-world case studies. Conversational agents can therefore be linked and identified from ordinary conversation, with no access to their weights or their hidden instructions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.