Structure Becomes Signal: Document Hierarchy as a Learned Behavioral Channel in LLM Evaluators
Abstract
Long-document LLMs process document content and the way it is organized. Yet document structure is typically treated as behaviorally neutral. We show that this assumption can fail: document hierarchy can become a learned behavioral signal. We study this phenomenon with Natural Document Structure poisoning (NDS), a controlled training-time intervention that associates a target heading-depth hierarchy with model behavior while preserving substantive content and source-specific heading vocabulary. With a 5% poisoning budget, NDS achieves 92.45 ± 0.63% target activation. Controlled interventions isolate the learned dependency on heading depth configuration: replacing heading text while preserving depth retains 91.82% activation, whereas changing only heading depths—while holding body text, heading strings, section order, and all non-heading tokens fixed—reduces activation to 6.92%. The effect persists across the tested target hierarchies, model families, document distributions, parsers, and a full-document question-answering task. These results show that, for long-document LLMs, how information is organized can become part of the behavioral input alongside what the document says. Finally, in an AutoResearch-style workflow, a simple outline constraint yields 72.50% structural compliance in generated manuscripts, of which 79.45% activate a fixed downstream evaluator. This reveals a broader threat: automated research systems can trigger learned behavioral dependencies through simple structural constraints, without requiring complex content manipulation or explicit malicious instructions in generated manuscripts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.