acceptodds
Under review as a conference paper at ICLR 2027

Private Federated Topic Discovery using On-Device Language Models

Abstract

Discovering recurring topics in user-generated text can guide improvements to AI and edge systems, but centralizing this text creates substantial privacy risks. We study distributed and privacy-preserving topic discovery, focusing on how user text is represented locally before private aggregation. We propose bounded semantic compression as a local representation method, whereby an on-device language model compresses each user's data into a bounded set of semantic terms. This addresses two limitations of existing the existing literature. First, it preserves more topic-relevant information under a fixed contribution bound, improving recovery after privacy noise is added. Second, it can name topics that are implied by the text but never stated directly, removing the requirement that a topic descriptor appear in the original text. We combine this representation with distributed private aggregation, providing user-level differential privacy for the released population topics without revealing individual user contributions. Across natural-language, synthetic, and source-code datasets, semantic compression improves topic recovery across private aggregation mechanisms, with the largest gains for multi-document users and topics that are implied rather than named.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.