CORPUS EVIDENCE AND FACTUAL RECALL CAUSAL TRACING WITH AUDITED LABELS
Abstract
Factual recall and corpus evidence are distinct observations: a model may answer a question even when the supporting subject–object pair is absent from a reference corpus. We compare activation-patching profiles under these operational labels on Pythia-2.8B and Pythia-6.9B. Across 2,216 traced fact–model pairs, recovery concentrates early in both networks. Among recallable facts, the groups have equal median peak depths, with no detected rank-distribution difference (p = 0.568/0.390). These reference-corpus labels do not establish training membership. We then compare independently restored single-layer rank-one interventions at relation-selected, fixed-middle, and control layers under equal relative update norms. On 100 held-out facts per model, selected and fixed-middle layers both achieve near-ceiling rewrite success. Their side effects differ: selected layers preserve other factual answers more often, by 3.3 and 7.0 percentage points. A post-hoc strength analysis reveals a target-success advantage at one-quarter of the calibrated norm: selected layers outperform the fixed-middle layer by 18 and 20 percentage points. Layer choice therefore interacts with intervention strength, distinguishing recovery location from a uniquely effective storage address. Repeating all dose cells with two additional full-pipeline split seeds preserves the quarter-strength advantage in both models, while showing that half-strength effects are small and model-specific. Transferring the selected layers to the non-deduplicated checkpoint twins yields quarter-strength advantages of 16 and 24 percentage points. A predeclared projected-gradient intervention saturates across sites and does not reproduce the selected–middle top-1 gap, bounding the layer-selection result to the rank-one update mechanism.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.