acceptodds
Under review as a conference paper at ICLR 2027

Poisoning Leaves a Trace: Post-Hoc Auditing and Recovery for Text Summarization

Abstract

Training-time data poisoning can induce persistent changes in abstractive summarization on ordinary inputs while preserving fluent outputs and conventional quality metrics, making compromise difficult to detect in open-ended generation. We study whether such trigger-free poisoning leaves diagnostic traces that can be exploited after fine-tuning. We identify two complementary signatures under different defender-access settings. When the fine-tuning corpus is available, the influence structure exploited by influence-targeted poisoning remains partially observable after fine-tuning; we use this signal to localize suspicious samples, validate their behavioral inconsistency, and selectively unlearn their effect. When the target model's fine-tuning data are unavailable, poisoned summarizers exhibit amplified sensitivity to controlled input perturbations. We capture this behavior through Sensitivity to Adversarial Perturbations (SAP), which measures perturbation-induced source-content exclusion for model auditing. We evaluate sentiment, toxicity, factual-distortion, and representational-bias poisoning, together with mixed-influence and mixed-objective attacks and position-shifted probing that stress the assumptions of the proposed defenses. Across nine architectures and six benchmark datasets, we detect poisoning under both defender-access settings while preserving summarization utility, and targeted unlearning achieves 84.6% average behavioral recovery at 10% contamination across poisoning objectives. These results show that trigger-free poisoning leaves measurable traces in training influence and model behavior, enabling post-hoc auditing and, when fine-tuning data are available, targeted recovery without full retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.