acceptodds
Under review as a conference paper at ICLR 2027

Structure Becomes Signal: Document Hierarchy as a Learned Behavioral Channel in LLM Evaluators

Abstract

Long-document LLMs process document content and the way it is organized. Yet document structure is typically treated as behaviorally neutral. We show that this assumption can fail: document hierarchy can become a learned behavioral signal. We study this phenomenon with Natural Document Structure poisoning (NDS), a controlled training-time intervention that associates a target heading-depth hierarchy with model behavior while preserving substantive content and source-specific heading vocabulary. With a 5% poisoning budget, NDS achieves 92.45 ± 0.63% target activation. Controlled interventions isolate the learned dependency on heading depth configuration: replacing heading text while preserving depth retains 91.82% activation, whereas changing only heading depths—while holding body text, heading strings, section order, and all non-heading tokens fixed—reduces activation to 6.92%. The effect persists across the tested target hierarchies, model families, document distributions, parsers, and a full-document question-answering task. These results show that, for long-document LLMs, how information is organized can become part of the behavioral input alongside what the document says. Finally, in an AutoResearch-style workflow, a simple outline constraint yields 72.50% structural compliance in generated manuscripts, of which 79.45% activate a fixed downstream evaluator. This reveals a broader threat: automated research systems can trigger learned behavioral dependencies through simple structural constraints, without requiring complex content manipulation or explicit malicious instructions in generated manuscripts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.