CKA's Blind Spot: Why Representation Analysis Misses Policy Change in RLVR, and How to Fix It
Abstract
Mechanistic interpretability asks whether representation-level tools can explain what RLVR does to a model. We find a sharp answer: a core such tool cannot. First, we show that the phenomenon motivating much of this line of work—RLVR "capability narrowing" (pass@k decreases for large k)—is a measurement artifact. Under a faithful multi-seed protocol it vanishes (Δpass@64 ≥ 0 across 4 seeds), and tracing it exposes three methodological pitfalls (evaluation protocol drift, fixed-dictionary SAE confound, spectral estimator noise) that plausibly affect 26–34 of 87 surveyed papers. As a frame for why policy-level changes barely move geometry, we prove a Strategy-Representation Gap theorem: policy-level (L1/L2) perturbations produce only bounded changes in representation geometry, with the bound vanishing under KL regularization; we validate it across scales, algorithms, and tasks (144× CKA discrimination; GRPO/DPO/Zero-RL mean CKA = 1.0000). Our central result is the CKA Blind Spot. Models whose representations CKA judges identical (CKA = 1.000000 to six decimals) produce measurably different outputs (KL = 0.143; 3–16% of top-1 positions disagree). Linear probing resolves the paradox: the strategy is present and linearly decodable (60–100% accuracy), so the gap is metric-specific, not information-theoretic. We prove CKA deviation is O(ε²) while probe sensitivity is O(ε)—a 1,000× gap at RLVR's ε 10⁻³ perturbation scale—and confirm it via weight interpolation (CKA flat at >0.9999 along the entire path while outputs phase-transition) and spectral decomposition (updates 2,500× smaller than base weights, yielding 1−CKA = Θ(ε²) 10⁻⁷, consistent with the saturated float32 reading). We fix the blind spot with NSARS, combining CKA with linear probing, which detects the gap at 5/8 layers vs. CKA's 1/8; causal mediation localizes strategy to early layers 0–6 (90.02% of divergence), living in a 5-dimensional subspace. A cross-paradigm boundary pins the mechanism: full-parameter RLVR leaves representations invariant (CKA = 1.0), while high-rank LoRA (7B, r=256) alters them dramatically (CKA = 0.059, erank −81%)—spectral smallness, not optimization, drives the blind spot. CKA should not be used alone: every CKA-based interpretability study should report linear-probe accuracy, and early layers deserve the attention now reserved for middle layers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.