acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Hidden Undesirable Behaviors from LLM Fine-Tuning via Consensus Distillation

Abstract

Fine-tuning can introduce unwanted behaviors through deliberate poisoning or unexpected generalization from seemingly benign data. Because these effects may be difficult to anticipate or detect, we seek robustness in the fine-tuning process itself. Our premise is that datasets independently collected for the same task can share the intended capability while differing in their unwanted effects. This creates an opportunity to learn from agreement across sources while reducing dependence on any individual source. We introduce Multi-Source Consensus Distillation (MSCD): separate teachers expose what each dataset teaches, a configurable consensus decoder regenerates training responses using agreement among those teachers, and a fresh student learns from the regenerated responses. We instantiate consensus with minimum and base-relative rules, then relax agreement to accommodate useful behaviors that only some teachers learn or that teachers express in different words. Experiments on controlled poisoning, subliminal learning, and emergent misalignment show suppression of unwanted behavior while retaining shared benefits. In a practical language-understanding task, MSCD retains intent-classification capability while reducing unsafe medical advice introduced during fine-tuning. Comparing distilled students with direct consensus decoding shows that distillation closely preserves the measured desirable and undesirable behavior rates while replacing multiple model evaluations per generated token with a single student evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.