Tracing Influence: Unboxing LLM Outputs with Sparse Latent Spaces
Abstract
A critical step toward the reliable use of large language models (LLMs) in high-stakes domains such as healthcare is to attribute their predictions to training data, akin to a medical case study. This requires fine-grained precision: pinpointing not only which training examples influence a decision, but also which parts of them are responsible. While influence functions offer a principled framework for this, prior token-level approaches are restricted to losses that decompose additively across tokens and distribute credit over a strongly correlated token basis. We introduce \method, a framework that attributes influence to internal features through a latent mediation approach and is defined for any differentiable objective. Our method attaches a sparse autoencoder (SAE) to an intermediate layer of a finetuned LLM to obtain a sparse, less correlated attribution basis, derives an activation-weighted first-order attribution of the training gradient, and scales it with derivative swapping, reducing the per-feature Jacobian-vector products to a single additional reverse-mode pass (13–18 faster). Token-level influence is then obtained by projecting latent attributions back to the input space via token activation patterns. Experiments on multiple-choice QA with Llama-3.2-1B and Qwen2.5-1.5B show that influence-selected features are more necessary and sufficient for the prediction than activation-, frequency-, or randomly-selected features, including under on-distribution replacement, and are better than a gradientactivation baseline in several settings while never significantly worse. The total feature influence preserves the magnitude ranking of training examples under direct influence (Spearman on Qwen2.5-1.5B), and reweighting the selected training data shifts predictions more than reweighting random data, supporting \method as a practical tool for model auditing.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.