acceptodds
Under review as a conference paper at ICLR 2027

NF-SAE: Testing an NMF-Inspired Sparse Autoencoder Across Language Models

Abstract

Neural networks store information as long lists of numbers called activations. Sparse autoencoders try to rewrite each list using only a few learned features. Inspired by nonnegative matrix factorization (NMF), a method that builds examples by adding nonnegative parts, we test an autoencoder whose main feature strengths cannot be negative. We call it NF-SAE. We train the same overall recipe on internal activations from three language models and compare it with a signed magnitude-TopK control. NF-SAE looks promising on Gemma, but the result does not carry over to Pythia or Qwen. BatchTopK has lower mean reconstruction and output errors in every setting where we test it. The clearest practical concern is that one fixed recipe behaves very differently across models of different sizes and scales; this study cannot show which design choice caused that difference. NF-SAE therefore remains an idea to test, not a generally better method. The main lesson is simple: success on one model should be checked with matched controls before making a broad claim.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.