acceptodds
Under review as a conference paper at ICLR 2027

SAEScientist-Bench: Evaluating LLM Agents on Concept-Driven SAE Feature Discovery

Abstract

Beyond improving model performance through training and data curation, enabling agents to study internal representations is a vital step toward closed-loop autonomous AI research and development (R&D). We introduce SAEScientist-Bench, a benchmark that evaluates whether large language model (LLM) agents can conduct interpretability research to discover concept features in sparse autoencoders (SAEs). Given a target concept, an agent investigates a 131K-feature dictionary for Gemma-2-9B-IT by designing contrastive probes, interpreting activation patterns across candidates, and submitting the feature that best captures the concept. To evaluate each submission, our benchmark provides held-out test suites containing positive texts, confusable negative texts, and prompts for text generation. Submissions are scored on three dimensions: prominence within the dictionary (Rank), selectivity in separating positive texts from negative texts (Activation), and the ability to steer generation toward the concept (Steering). We find that agents reliably isolate features that distinguish texts, reaching 92.9 Activation against a reference feature's 98.9, but these features struggle to steer generation, reaching at most 31.5 Steering against 57.8. This gap stems from candidate selection: agents usually uncover and test strong candidates during their investigation, yet end up choosing features that appear selective on probe texts but fail during generation. Moreover, language concepts prove the hardest category due to confusion across related languages. These findings suggest that conducting rigorous interpretability research remains a critical capability that agents have yet to master. Code is available at https://anonymous.4open.science/status/SAEScientist.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.