When Reasons Begin to Tilt: Reasoning Prior Implantation in Large Language Models
Abstract
As LLMs are increasingly used to mediate high-stakes decisions, their behavior depends not only on factual knowledge but also on how they weigh competing reasons, constraints, and justifications. However, open supply chains make LLMs vulnerable to backdoor attacks, and existing threats heavily rely on a trigger-to-behavior paradigm that forces attacker-specified surface outputs. These shallow interventions fail to generalize to nuanced ethical or logical judgments and are easily erased by downstream adaptations. In this paper, we introduce ThoughtSeal, the first backdoor threat that shifts the threat target from an LLM's surface actions to the deep mindset from which it reasons. Instead of dictating what a model should say, ThoughtSeal implants a reasoning prior that corrupts how it thinks, judges and decides. To achieve this goal, we propose ThoughtSeal Preference Distillation (TPD). TPD canonicalizes an abstract adversarial mindset and internalizes it via on-policy preference distillation over the model's own competing rationales, alongside context-agnostic persistence optimization. Extensive evaluations demonstrate that ThoughtSeal exhibits capabilities beyond surface-level backdoors: it activates via natural concepts without artificial triggers, seamlessly transfers across unseen domains, robustly survives downstream adaptations, and faithfully preserves general task utility. Ultimately, our work exposes a profound vulnerability: a compromised model can remain entirely fluent and helpful, yet silently operate as a hidden arbiter across diverse downstream decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.