Blind to the Latent Truth: LLM Endognostics and the Schizognosis of Minerva-7B
Abstract
Black-box evaluation through generated text relies on an unverified premise: that surface tokens faithfully reflect hidden-layer representations. Preference alignment rewards output compliance rather than epistemic accuracy, which may push the decoding policy to discard internal factual states. We call this representational split "schizognosis": hidden activations encode truth or risk while the policy emits compliant or false responses. We evaluate the phenomenon in Minerva-7B-Instruct-v1.0 (32 layers, 2.48 trillion pre-training tokens) using white-box endognostic projections. Across 124 contrastive minimal prompt pairs (12 legal, administrative and professional risk domains), behavior is identical on 63.7% of pairs (79/124; 95% CI [55.0%, 71.6%]): dual compliance on 59 (47.6%), prudential refusal of both members on 20 (16.1%). Projecting residual-stream activations onto vocabulary space via a Jacobian lens isolates a Contrastive Endognostic Margin (CEM); under Holm-Bonferroni correction (FWER 0.05) two categories remain significant: Sycophancy Signal (, Cohen's ) and Confidentiality-Disclosure Risk (, ). The model accepts a presupposed falsehood on 18 of 25 facts when the false premise sits inside a secondary question, against 1–4 when it is asserted and then questioned (Cochran's , ). The correct entity token stays readable on 25 of the 29 measurable errors (86.2%), inside a late intermediate "verbalizable workspace" (layers 19–29, 61–94% of depth). Ablating the direction of the planted falsehood there restores ground-truth outputs in 11 of 25 suppressed cases (44.0%, 95% CI [26.7%, 62.9%]), against none for a random direction of equal norm (exact McNemar ), an unrelated token, or the true token; outcomes are labelled by two blind human annotators (). An out-of-sample linear probe reaches 77.0% accuracy at layer 10 (32% of depth) and 61% inside the band; ablating its in-band directions recovers nothing (0 of 25) and changes only 3 outputs. Linear decodability therefore does not imply causal control over verbalization. The Jacobian transport is what reads the suppressed truth (25 against 21 detections for the plain unembedding), while the causal lever is the vocabulary direction of the false token: evaluating aligned models on their answers alone misses the common case, not a corner of it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.