Sparse-Circuit Recoverability Across Grokking and Decision-Matched Realisations
Abstract
Does the circuit-recovery advantage associated with grokking depend on a network's original weights, or does it recur in a new realisation of its decisions? We test this by combining longitudinal analysis of behaviourally sufficient component masks with exact, full-domain decision matching. In a five-seed modular-addition study, greedy recovery at 99% decision agreement retains 510–515 of 516 components at delayed checkpoints, compared with 65–146 after stable generalisation. Targeted non-monotone and distribution-sensitive follow-ups retain a large continuous gap. We then train independently initialised students to reproduce each teacher's complete decision table, including its errors, without constraining logits or internal representations. Under greedy recovery against each model's own centred logits, the stable-minus-pre contrast is negative for teachers and student-cell medians in all fourteen paired addition seeds and, descriptively, all eight paired multiplication seeds. Mean reductions are 21.3 and 12.0 percentage points for addition teachers and students, respectively, and 8.8 points for both on multiplication. Absolute endpoints vary across decision-matched realisations, while the executed hard-concrete policy returns predominantly near-intact masks. Under the tested retraining and greedy-recovery procedures, the phase ordering therefore recurs without the original teacher weights, while its magnitude and absolute endpoints change.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.