acceptodds
Under review as a conference paper at ICLR 2027

Black-Box Forensics for Conversational LLM Agents

Abstract

LLM-powered scams are proliferating, API providers silently change the models behind their endpoints, and jailbreaks are increasingly model-specific. Each of these calls for black-box forensics of conversational LLM agents: learning what powers an agent from nothing but ordinary, non-adversarial conversation with it. We study two capabilities. Attribution identifies which of a known set of base models or system prompts powers an agent, and we treat it as classification: a classifier trained on stylistic and n-gram features identifies the base model with 98% accuracy from a few turns of conversation. System prompts can be attributed the same way, but only by collecting labeled conversations for every candidate prompt, and the prompts deployed in the wild are too many and change too often for that. Fingerprinting instead decides whether two conversations come from the same agent, and so needs no training data from the agents it is applied to. Our cross-encoder scores a pair of conversations zero-shot: on system prompts never seen in training, it reaches an AUC of 0.768 from a single pair of conversations, and 0.943 when 50 conversations per agent are aggregated. It remains useful when every reply is paraphrased by another model, and it gives positive results on real-world case studies. Conversational agents can therefore be linked and identified from ordinary conversation, with no access to their weights or their hidden instructions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.