SAEScientist-Bench: Evaluating LLM Agents on Concept-Driven SAE Feature Discovery
Abstract
Beyond improving model performance through training and data curation, enabling agents to study internal representations is a vital step toward closed-loop autonomous AI research and development (R&D). We introduce SAEScientist-Bench, a benchmark that evaluates whether large language model (LLM) agents can conduct interpretability research to discover concept features in sparse autoencoders (SAEs). Given a target concept, an agent investigates a 131K-feature dictionary for Gemma-2-9B-IT by designing contrastive probes, interpreting activation patterns across candidates, and submitting the feature that best captures the concept. To evaluate each submission, our benchmark provides held-out test suites containing positive texts, confusable negative texts, and prompts for text generation. Submissions are scored on three dimensions: prominence within the dictionary (Rank), selectivity in separating positive texts from negative texts (Activation), and the ability to steer generation toward the concept (Steering). We find that agents reliably isolate features that distinguish texts, reaching 92.9 Activation against a reference feature's 98.9, but these features struggle to steer generation, reaching at most 31.5 Steering against 57.8. This gap stems from candidate selection: agents usually uncover and test strong candidates during their investigation, yet end up choosing features that appear selective on probe texts but fail during generation. Moreover, language concepts prove the hardest category due to confusion across related languages. These findings suggest that conducting rigorous interpretability research remains a critical capability that agents have yet to master. Code is available at https://anonymous.4open.science/status/SAEScientist.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.