acceptodds
Under review as a conference paper at ICLR 2027

VERITAS: VERified Interpretations of Transformer ActivationS

Abstract

Activation interpretability faces a fundamental credibility gap: a learned reader can reconstruct a hidden state through a bottleneck that is semantically wrong. Reconstruction constrains information flow; it does not fix meaning. We introduce VERITAS, which closes this gap by replacing self-authored semantics with executable semantics. Before any activation is read, an external executor runs a finite set of candidate actions, normalizes their observed successor states, and exposes the resulting codebook to the reader. The reader may select a candidate identity but cannot redefine what that identity denotes—denotation is owned by the verifier, not the interpreter. We instantiate VERITAS in Lean theorem proving with Goedel-Prover-V2-8B across 1,666 proof transitions. Four complementary measurements establish the interface's properties. Externally fixed candidate identity improves held-out activation prediction by ΔFVE_id = 0.0454 beyond a decoder that already sees the full candidate set (95% CI [0.0291, 0.0623]). An activation-conditioned pointer reaches 80.9% candidate-preference accuracy versus 73.6% for the strongest non-activation baseline. Candidate-conditioned reconstructions steer the frozen model toward their target candidate with a 0.114 argmax-match advantage over matched arbitrary-code reconstructions (95% CI [0.077, 0.158]). And while state-aware attackers readily recover an injected nuisance signal from an equal-cardinality free code, the grounded pointer yields near-zero incremental leakage across a 30× attacker-capacity range. The core reconstruction and decoding signals replicate in the Qwen3 parent model under the same candidate environment. VERITAS points to a broader design principle: let an executable system own the semantics of the latent channel, then probe how neural states predict, reconstruct through, and steer via it. Lean is the first testbed; the same interface extends naturally to sandboxes, unit tests, simulators, and typed tool effects—turning activation interpretability into an actionable foundation for white-box AI control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.