acceptodds
Under review as a conference paper at ICLR 2027

CBM-Scientist: Agent-driven Automated Concept Bottleneck Model Design

Abstract

Automated research agents hold significant promise for accelerating machine learning research, yet their utility in designing inherently interpretable neural networks remains understudied. We introduce CBM-Scientist, an agent-driven framework for automated concept bottleneck model (CBM) design. Agents iteratively propose, implement, and evaluate model and training changes, using experimental feedback to guide subsequent proposals. We study four proposal strategies across image classification, text generation, question answering, and image generation. Across these settings, agents develop recipes that improve multiple evaluated objectives simultaneously. For example, on AGNews, a loss-constrained recipe improves concept F1 by 15.10 % and steerability by 44.27 % while reducing perplexity from 96.23 to 43.17. Analysis of sprint trajectories shows how agents combine modifications, revisit unsuccessful ideas, and refine training procedures when improvements in one objective do not transfer to others. Together, these findings demonstrate the potential of autoresearch for concept bottleneck model development and highlight the importance of evaluating both task performance and concept use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.