Utility, Sensitive-Content Exclusion and Privacy in Graph-Based Summarization
Abstract
Summarizing sensitive documents requires useful content, control over sensitive source spans, and a precise privacy guarantee. We develop a framework that separates these requirements and connects each to testable evidence in a source-backed graph summariser. We prove selection privacy under fixed public support and bounded score sensitivity, and construct counterexamples showing why stable filtering alone cannot extend this guarantee to private extractive outputs. We establish source-level exclusion under detector coverage and provenance conditions, and derive a tight objective-headroom bound on selection gains over uniform feasible summaries. Across eight corpora and 355,200 selection records, all 135 Holm-significant exponential-versus-frontier effects are positive. Controlled disease-label interventions isolate the effect of protection; evaluations on 100 biomedical abstracts and 48 clinical cases quantify the gap between excluding detected spans and covering every sensitive span. Larger protection radii reduce observed exposure at the cost of more empty releases. A 240-cell provider-adapted generation study complements these results with coverage gains and explicit bounds for its 11 failed cells. Together, the analysis and experiments identify which evidence supports each release claim and which design choices separate content protection from differential privacy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.