VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
Abstract
Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly shifting toward *vibe protein design*, where a single model is expected to generate novel sequences, engineer existing proteins, and reason about protein characteristics through flexible natural-language constraints. However, existing benchmarks typically evaluate isolated aspects of protein design or assume predefined structured inputs, making them ill-suited to assess this broad, open-ended setting. To address this gap, we present Vibe Protein Design Benchmark (**VibeProteinBench**), a language-interfaced benchmark that evaluates whether a single model can operate across three complementary stages of a computational protein-design workflow: **recognition**, **engineering**, and **generation**. Each stage is grounded in expert-curated mechanistic rationales and multi-faceted in silico validation, to computationally verify whether model outputs are biologically plausible. Across diverse general-purpose and domain-specialized LLMs, evaluated with and without tool access, no model performs strongly across all three stages at once. Beyond measuring end-task success, **VibeProteinBench** exposes where current systems break: we identify recurring failure patterns and characterize how tool-augmented agents select tool arguments. Together, these findings show that vibe protein design remains a substantial open challenge.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.