Frequency, Direction, and Feature Sharing: Dissecting Co-activation in Sparse Autoencoders
Abstract
Co-activation in sparse autoencoders (SAEs) is often used to infer concept hierarchies, but shared features and parent–child direction require different evidence. We show that on non-tied pairs, reciprocal coverage orients concepts solely by support-set size. In a complete noiseless chain model with an unknown full-rank dictionary, every ordering of L nonzero states admits an exact representation, and an L₀ penalty favors frequent states at shallow depths rather than identifying semantic order. On 5,572 WordNet pairs with Llama-3.1-8B/JumpReLU, direction accuracy is 79.1% when the child-to-parent corpus-frequency ratio is below one and 25.3% otherwise. In a separate matched-control analysis, a lexical-frequency baseline scores 66.1% versus 61.1% for the support rule, while hierarchical pairs show greater excess overlap than frequency-matched random controls (Cohen's d=0.91). Replacing JumpReLU with TopK at fixed encoder weights and approximately matched mean activation counts reduces the frequency-regime gap from 53.8 to 6.4 points. Our results distinguish frequency-associated orientation from semantic sharing and give sufficient conditions for latent-order recovery within the chain model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.