PARSE: An Autonomous Algorithm for Mechanistic Interpretability Based on Dynamic, Adaptive Sparse Autoencoders
Abstract
The resolution of semantic superposition in Large Language Models (LLMs) currently relies on sparse autoencoders (SAEs) optimized via static regularization and heuristic warm-up schedules. These static schedules often force concepts into entangled representations and induce semantic hallucination. We propose that representation collapse is not an inherent flaw of SAEs, but an artifact of static latent expansion which lacks quantitative feedback on the condition of the feature covariance matrix. In this paper we present a novel autonomous algorithm that eliminates heuristic warm-up schedules by allowing sparsity to emerge from the dynamic linear algebraic properties of the covariance matrix under the action of the optimizer. PARSE (Parabolic Adaptive Relaxation for Sparse Encoding) continuously monitors the geometric strain of the latent manifold via a matrix-free Krylov approximation of the condition number. When local semantic cross-talk exceeds the LLM's inherent noise floor, the algorithm dynamically spawns orthogonal vectors strictly aligned with the principal eigenvectors of the local residual error matrix. By temporarily substituting the penalty with a soft orthogonalization penalty, PARSE absorbs dense concepts and safely recycles dead latents without inducing representation shock. This localized expansion reduces the computational profile from quadratic to linear scaling, eliminating GPU bottlenecks and reducing overhead by over 90%. We validate PARSE on the 4096D residual stream of Llama-3(8B) across tokens and demonstrate that complex semantic superposition can be isolated with low error under the constraint of strict numerical orthogonality. We further demonstrate that feature death and representation collapse are distinct phenomena. PARSE thus establishes a scalable, autonomous foundation for mechanistic interpretability and the translation of emergent frontier model concepts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.