acceptodds
Under review as a conference paper at ICLR 2027

True Quotes, False Impressions: A Benchmark for Adversarial Extractive Summarization on Long-Document Classification

Abstract

LLM judges often decide from evidence selected by an upstream system. We ask whether they can be misled through verbatim sentence selection alone. We introduce a long-document classification benchmark spanning sentiment reviews, partisan news, and legal fact patterns. A selection-only extractor may choose source sentences but cannot rewrite, reorder, or fabricate them, and the judge sees only the selected evidence. On documents selected by a lexical heuristic that favors mixed evidence, adversarial selection changes – of the decisions that the same judge makes correctly from the full document, pooled over six extractors and three judges. A blind LLM audit codes every excerpt in of audited successful selections as either unchanged or weakened, but not reversed, by its context. Full-document access removes most selection-induced flips on every dataset. The results show that requiring evidence to be copied exactly from the source does not prevent strategically selected excerpts from misleading the judge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.