Before The Attacker Acts: Evaluation Floors for LLM Deanonymisation
Abstract
Deanonymisation scores for language-model attackers come with no reference level, so a reader cannot tell how much of a score is inference. We give two levels. Below is a floor: the best score of null attackers that read no post and claim no more than the task’s declared volume v (the seed, random guessing at v, the modal value, abstention), fixed before the attacker runs, with a matched volume contrast added afterwards; every lift is read under one declared paired t-test. Above is a labelled-evidence ceiling: the recall of an attacker merging accounts only on specifics the answer key records, which at the planned number of personas gives a pre-run screen, saying from the key alone whether any such attacker could produce a detectable lift. Under SynthPAI’s own scorer a constant guess fitted out of sample scores 0.275 against GPT-4’s 0.780; all 18 models clear it overall, but after Holm correction over 144 model-by-attribute cells 65 clear, 1 is at or below it and 78 are undecided, none on age or income. On one of our corpora the screen finds that no attacker merging only on labelled evidence could be told from random on its 20 personas, and the agent’s cross-account recall of 0.121 sits below the 0.137 random guessing at v = 3 handles per seed is expected to reach, though the agent claimed one handle in all. Report the null row, the declared test and the pre-run screen beside the score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.