acceptodds
Under review as a conference paper at ICLR 2027

Influential but Misaligned: LLM Evidence Drives Predictions Yet Misses What Institutions Classified

Abstract

High-reliability security document classification requires knowing not only how accurately language models predict but also what evidence supports their decisions. Prior work has evaluated the predictive influence of model rationales and their agreement with human annotations, but research annotations alone cannot establish alignment with records of actual classification practice. We use portion markings in declassified documents released through the CIA Electronic Reading Room as a comparison reference. Portion markings record classification levels for document spans; where the highest-level rule applies, the highest-classified portions are directly linked to the document label. Using texts with classification markings removed, we evaluate the predictive influence and selection consistency of self-rationales from three open-weight language models and compare them with institutional markings in the portion-recovered subset. In a stratified sample, deleting the top self-rationale changed predictions 10.5–12.4 percentage points more often than length-matched random deletion. Rationale selection was stable under identical inputs but varied with prompt wording and model choice. In the portion-recovered subset, model rationales overlapped institutional portions less than a within-document random-position baseline (lift 0.65–0.75), and we found no evidence that deleting classified portions changed predictions more often than amount-matched random deletion. This divergence also left a positional and lexical signature. Llama rationales appeared near the beginning of documents more often than institutional portions and contained classification and secrecy vocabulary 26.6 percentage points more often in a sentence comparison matched for position and length. This pattern is consistent with a preference for sentences that overtly signal classification. The predictive influence of model rationales and their alignment with institutional span-level classification records therefore require separate evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.