acceptodds
Under review as a conference paper at ICLR 2027

From Clues to Sensitive Disclosure: Measuring Multi-Turn Privacy Risk and Preventing Privacy Leakage in LLMs

Abstract

Multi-turn LLM assistants can expose sensitive information even when no sin- gle historical message states it: individually weak clues can jointly support a new person-level conclusion. Existing privacy checks rarely explain which ear- lier interactions support such a conclusion or enforce protection before the first external release. We introduce PRIVACT, a framework that distinguishes source- explicit disclosures from Novel Sensitive Claims (NSCs), measures historical sup- port through a temporally ordered Privacy Trajectory, and buffers candidate output while forecasting disclosure risk and selecting a pre-release action. On SynthPAI, the observed NSC rate for Qwen2.5-7B-Instruct rises from 3.92% with shallow history to 25.92% with full history. On 1,000 TOP-Bench cases, historical attribu- tion achieves 0.803 Exact Support. In the matched SynthPAI protection compari- son, PRIVACT achieves 100% detection precision and 79.31% prevention; across evaluated generators, prevention reaches 86.14% on Mistral-7B, with a measured end-to-end latency increase as low as 1.32% on Qwen2.5-7B. These results show why multi-turn privacy protection must account for both how a claim forms and when it becomes externally observable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.