Challenges in Evaluating Contextual Integrity for Large Language Models
Abstract
Modern large language models are powerful tools that can benefit consumers across many domains and applications, but their ability to interact with and use personal information creates a risk of privacy violations when information is improperly shared. This has led to the development of privacy benchmarks for LLMs based on contextual integrity, a framework for evaluating privacy in terms of information flow within specific contexts. We identify key limitations of these benchmarks, such as missing data/code, reproducibility, and incomparable results, that make it challenging to understand how evaluative claims for contextual integrity are interpreted. We further present a critical analysis of three open-source benchmarks with provided data, focusing on areas like data quality, variance, and metric choice. These findings provide additional context for existing benchmarks, informing both practitioners who use them and future benchmark developers about opportunities to improve how benchmarks can more accurately assess contextual integrity in LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.