Under review as a conference paper at ICLR 2027
CALLIOPE: Private Insights into AI Use, without Embeddings
Abstract
We propose _CALLIOPE_, a top-down framework for discovering topic hierarchies in large corpora of unstructured data (e.g., LLM conversations), while supporting differential privacy (DP). Unlike traditional and even recent LLM-augmented approaches that ultimately rely on clustering vector embeddings, our framework leverages LLM capabilities directly for the entire pipeline. Experiments demonstrate that this approach provides better utility over existing (embedding-based) methods in the DP setting, especially in the high privacy regime. We evaluate the quality of _CALLIOPE_ outputs using a variety of traditionally studied clustering metrics, as well as some novel utility metrics that we introduce.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.