Moonshine: A Protocol for Distillation in Unsafe Model Ecosystems
Abstract
Organizations without access to U.S. frontier intelligence increasingly rely on open weight models for agentic cyber-defense. Because the provenance of open weight models cannot currently be verified, these organizations are exposed to the sleeper agent supply chain risk; their cyber-defense model may appear safe, but actually contain a backdoor that can trigger unsafe behavior after deployment. To address this, we introduce Moonshine, a protocol for distillation in unsafe model ecosystems, where neither the teacher model nor the actor performing the distillation can be trusted. Moonshine combines an attested training pipeline, which reduces trust in the distillation actor, with hard sequence and on policy distillation, which both restrict the logit channel used to smuggle backdoors through distillation. By implementing Moonshine, a distillation actor can attest to training a cyber-defense model using distillation methods that are empirically shown to remove backdoors, thereby mitigating the sleeper agent supply chain risk. We provide a proof sketch of protocol soundness and empirical results demonstrating that hard sequence and on policy distillaton induce capability transfer without causing backdoor transfer across a wide variety of model and trigger types, including triggers explicitly designed to persist through distillation. We conclude by arguing that adoption of Moonshine would differentially accelerate agentic-cyber defense without accelerating agentic cyber-offense.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.