acceptodds
Under review as a conference paper at ICLR 2027

BioCausalX: An Environment for Causal Discovery in Biological Systems

Abstract

Large language model (LLM) agents combine extensive prior knowledge with limited experimental evidence when investigating biological systems. Evaluating and training these agents requires repeatable cycles of intervention, observation, and hypothesis revision. Biological experiments are often slow and costly; simulated environments offer repeatable feedback but also require scientific descriptions and interventions tied to documented mechanisms. We introduce BioCausalX, an environment for causal discovery across 43 executable biological systems and 168 source-supported intervention axes. Variable descriptions and permitted operations are reconciled with publications and database records. Each operation maps to an executable mechanism hidden from the agent, with supporting evidence retained for audit. Across evaluated LLMs, biological context supports recovery, while alignment with the underlying system shapes the benefits of experimentation. Within tested budgets, feedback partly compensates for missing descriptions but does not reliably overcome misleading ones. Larger experimental allowances can remain underused without improving recovery, identifying limitations in both resource allocation and evidence integration. GRPO training improves recovery on held-out systems with and without biological context. Together, these results demonstrate that BioCausalX supports reproducible evaluation and training of scientific agents that combine biological knowledge with experimental feedback.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.