acceptodds
Under review as a conference paper at ICLR 2027

Which Features Matter? Sparse Latent Intervention for Prediction-Preserving Explanations

Abstract

Deep neural networks achieve strong predictive performance across a broad range of domains, yet their internal representations remain largely opaque: for a given prediction, it is generally unclear which internal features the model relies on, and whether they correspond to meaningful structure in the underlying domain. Sparse autoencoders (SAEs) decompose hidden representations into interpretable features, but reveal only what a layer represents, not which features a prediction actually depends on, while intervention-based methods typically assess features one at a time and overlook their interactions. We present Sparse Latent Intervention (SLI), a mechanistic interpretability framework that identifies, for a frozen model and a single input, a sparse set of SAE features that closely reproduces the model’s prediction. SLI casts this inverse problem as a sparse selection problem: stochastic gates over the entire feature dictionary are jointly optimised under an expected- penalty to preserve the prediction, while the counterfactual representation induced by the selected features is obtained in closed form through the frozen decoder, making the procedure end-to-end differentiable. We demonstrate SLI on a Transformer-based transcriptomic foundation model for two clinical prediction tasks, treatment response in inflammatory bowel disease and synovial endotype in rheumatoid arthritis, where curated domain knowledge enables independent validation, and compare it against established attribution methods. On both tasks, SLI recovers compact feature sets that closely reproduce the model’s predictions, recur across patients, and align with domain-relevant structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.