EvoCircuit: Active Learning for Protein Optimization via Mechanistic Interpretability
Abstract
Protein language models (pLMs) provide rich representations for protein engineering, but active optimization must adapt these representations to target-specific objectives from only a small number of experimental measurements. Existing protein optimization approaches improve candidate selection through activity prediction or search over learned fitness landscapes, but do not explicitly exploit task-relevant sparse representations within the underlying pLM. We introduce EvoCircuit, an active-learning framework informed by mechanistic interpretability that uses sparse features extracted from cross-layer transcoders of ESM-2. At each round, EvoCircuit identifies task-relevant features through associations between feature activations and experimental activity in measured variants. These features induce a task-adaptive similarity over candidate sequences, guiding batch acquisition alongside predictor-based ranking. Across retrospective active-learning simulations on single-mutant protein fitness landscapes, EvoCircuit improves the discovery of high-activity variants over predictor-based selection under matched assay budgets, with effects varying across landscapes and predictors. Ablations examine the contributions of task-adaptive feature selection and acquisition design, including comparisons with embedding-based and random-feature controls. Experiments on a combinatorial multi-mutant landscape extend the evaluation to predictor-based and reward-guided search baselines, revealing differences between hit discovery and peak-fitness optimization. Our results suggest that sparse internal representations of pretrained pLMs can complement activity prediction as inductive signals for data-efficient protein optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.