Trellis: Agent-Orchestrated Design and Evolution of Biological Sequence-Based Property Predictors
Abstract
Predicting biological properties such as expression, immunogenicity, stability, and binding affinity of proteins is a critical step in many drug discovery pipelines. New experimental datasets and protein sequence representations create opportunities to improve these predictors, but updating them requires repeated engineering of features, architectures, and training procedures. To address this, we introduce TRELLIS, an agent-orchestrated framework for developing biological sequence based property predictors. Large language model (LLM) agents write and execute code in isolated environments, and a Monte Carlo Tree Search explores diverse combinations of sequence representations and model architectures. Across FLIP2, PeptiVerse, and major histocompatibility complex (MHC) class I pathway tasks, a multi turn tool calling agent produced predictors with scores above the listed literature references on eight tasks. Tree search improved test performance over the best Qwen3.8-27B baseline on seven of eight selected tasks. We also evaluated sequential updates of an MHC class I elution predictor using Immune Epitope Database (IEDB) release-year holdouts. We release our code, which accepts any labeled sequence-property dataset without task-specific modification at https: //anonymous.4open.science/r/trellis-B12C.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.