acceptodds
Under review as a conference paper at ICLR 2027

Echo Poisoning: Feedback-Amplified Attacks in On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student language model on its own responses using teacher feedback. Prior distillation attacks transfer malicious behavior from a compromised teacher but overlook that poisoned feedback can make the student create its own poisoning opportunities. We propose Echo Poisoning, which edits only the teacher probabilities returned during OPD. It reinforces trigger phrases the student begins on its own and teaches attacker-chosen text after each completed trigger, so more frequent triggers bring more poisoned supervision. We call this process feedback amplification. Across three teacher-student pairs from three model families on math and code tasks, Echo Poisoning at least triples the attack success of payload-only poisoning, and accuracy loss is 1.8 to 8.0 points. Constraining trigger growth removes this advantage. Existing defenses leave part of the attack in place: divergence masking blocks payload delivery but not trigger growth, and other defenses without trigger knowledge leave 14.4% to 31.9% attack success. We therefore propose Sentinel, which replaces suspicious teacher targets with predictions from a frozen copy of the initial student. Sentinel blocks all observed non-adaptive delivery, returns trigger frequency close to the clean level, and keeps attack success at 2.0% under an adaptive attack that evades its detector. Code and configuration files are available at https://anonymous.4open.science/r/echo-poisoning-sentinel-34B7/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.