acceptodds
Under review as a conference paper at ICLR 2027

Chat Templates Are a Hidden Instruction Channel: Inference-Time Backdoors in LLM-Based Systems

Abstract

Open-weight language models are distributed as self-contained artifacts bundling weights with an executable chat template, a program (usually Jinja2) invoked at every inference call to format inputs before the model sees them. We identify this component as an important but overlooked attack surface. An adversary who redistributes a model with a modified template implants an inference-time backdoor without touching weights, training data, or deployment infrastructure; the only capability required is publishing the model's file. We evaluate the attack across three tiers. At the model level, triggered backdoors cut grounded factual-QA accuracy from 90% to 15% on average and induce attacker-controlled URL emission with over 80% success, while benign inputs show no measurable degradation, across eighteen models, seven families, and four inference engines. At the agent level, template backdoors hijack tool use across two benchmarks spanning 3,852 task runs, remaining effective under the tested input-level defenses while staying dormant absent the trigger. At the system level, a single poisoned artifact compromises a real-world agentic deployment and propagates supply-chain code poisoning downstream. The artifacts evade all automated scans on the largest model hub and remain unreachable by well-known defenses. Together, these results expose chat templates as an overlooked supply-chain trust boundary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.