acceptodds
Under review as a conference paper at ICLR 2027

Bringing Model Internals into Expert Annotation of Legal Reasoning

Abstract

Large language models (LLMs) increasingly perform complex professional work, creating a need to examine the reasoning behind their decisions. In legal tasks, written rationales alone may provide an incomplete basis for expert oversight. We introduce an activation-informed expert annotation framework in which legal experts inspect generated reasoning alongside natural-language explanations derived from model activations. Three experts contributed 34 hours of review, producing 354 span-level legal annotations across 15 conversational and agentic trajectories, together with explanation-helpfulness ratings for eight trajectories. To supply explanations across models, we propose Cross-LLM Activation Surrogacy: an accessible surrogate model reads a target model’s trajectory, and a pretrained Natural Language Autoencoder (NLA; Fraser-Taliente et al., 2026) describes the surrogate’s resulting activations. This approach requires neither access to the target’s activations nor training a dedicated NLA for each target. Experts judged NLA-based explanations helpful in 41.4% of their ratings. In classification of previously identified legal errors, aligned surrogate explanations improve macro-F1 from 0.257 to 0.324 relative to providing no explanations, with gains also observed over text-only and shuffled-explanation controls. Cross-model comparisons show stronger semantic correspondence between explanations within model families than across families, while reconstruction similarity alone does not reliably track this correspondence. Together, these results establish a practical framework for activation-informed expert annotation, with initial evidence that aligned surrogate explanations support classification of expert-identified legal errors. Code and dataset are available at: https://anonymous.4open.science/r/tr3D8F.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.