When does redundant supervision resist preference contagion? Path convergence and candidate-dependent scoring
Abstract
Peer-scored preference learning allows a compromised language-model agent to influence learners beyond its direct supervision. Redundant supervision from multiple teachers offers an intuitive defense, but its effectiveness is unclear when teachers themselves repeatedly update. We study this problem in multi-agent Direct Preference Optimization (DPO), focusing on two factors missed by static teacher counts: propagation through the supervision graph and candidate-dependent scoring. Local redundancy can fail under path convergence: a two-teacher ring remains clean, whereas in a diamond a single poisoned source infects two relays whose feedback converges on a learner with the same in-degree, with the pattern reproduced across Qwen2.5-1.5B and Llama-3.1-8B. A permanently clean third teacher delays, but does not prevent, infection. Holding teacher weights, prompts, and clean candidates fixed, the clean teacher’s mean clean-over-poison margin drops from 4.617 to 2.092 nats/token when moving from constructed poisoned candidates to candidates generated by a scoring teacher. We then preregister predictions for four held-out topologies using a closure over contaminated-teacher fractions. The closure matches all four outcomes, compared with one for static in-degree and three for source reachability, and an uncalibrated strict-majority closure matches as well. Redundant supervision is therefore not determined by teacher count alone. Its safety depends on how contamination propagates through the supervision graph and on the candidates over which supervision is applied.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.