True Quotes, False Impressions: A Benchmark for Adversarial Extractive Summarization on Long-Document Classification
Abstract
LLM judges often decide from evidence selected by an upstream system. We ask whether they can be misled through verbatim sentence selection alone. We introduce a long-document classification benchmark spanning sentiment reviews, partisan news, and legal fact patterns. A selection-only extractor may choose source sentences but cannot rewrite, reorder, or fabricate them, and the judge sees only the selected evidence. On documents selected by a lexical heuristic that favors mixed evidence, adversarial selection changes – of the decisions that the same judge makes correctly from the full document, pooled over six extractors and three judges. A blind LLM audit codes every excerpt in of audited successful selections as either unchanged or weakened, but not reversed, by its context. Full-document access removes most selection-induced flips on every dataset. The results show that requiring evidence to be copied exactly from the source does not prevent strategically selected excerpts from misleading the judge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.