acceptodds
Under review as a conference paper at ICLR 2027

Computational Provenance in Pretrained Language Models

Abstract

Verifying why a language model produced a response requires knowing which internal states contributed to its generation, yet the resulting output alone does not reveal which activations were responsible. We extend computational provenance—evidence linking an output to the causally relevant internal computation that produced it—from purpose-built neural architectures to pretrained language models. Our construction forms a discrete state from frozen activations, makes downstream computation depend on it, authenticates a trusted observation of the state used, and carries that authorised state into constrained generated output from which it can later be recovered. In separate sealed evaluations, Gemma 3 1B achieved 98.34% state-prediction accuracy on 4,096 examples and Qwen3-0.6B achieved 95.70% on 256 examples; all accepted executions passed the registered causal, authentication, and recovery checks, while negative controls produced the required rejection or abstention. Separately, a state identified in unmodified Gemma activations passed a causal test, with 32/32 answer-preserving and 30/32 answer-changing interventions, but did not reach a valid test of recovery from ordinary generated text. In a developmental study, a less constrained text-generation variant recovered the authenticated state in 27/32 cases at top-1 and 32/32 at top-4, but did not meet all reliability requirements. These results establish end-to-end bounded computational provenance in two pretrained language models through an engineered finite-state pathway, and show partial progress toward naturally occurring states and less constrained text.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.