When Tools Dominate, When Strategies Matter: A Causal Study of Agentic Computational Biology
Abstract
Agentic AI increasingly automates computational biology workflows, yet performance varies widely and inconsistently across benchmarks, and observational studies cannot separate tool heterogeneity from planning-strategy differences. We introduce Dissect-Audit, a causal-study framework that isolates the two via controlled factorial experiments on single-cell RNA-seq analysis, using a shared registry of validated operations to independently vary the tool layer and planning strategy under do-calculus logic. Across our 23-task core benchmark, tool-layer unification alone eliminates inter-strategy variance (Wilcoxon ) for all eight frontier LLM planners tested: every planner gains quality and sheds run-to-run variance ( to , mean ; variance reduced 42–100%), and planners sharply separated under native tools become statistically indistinguishable (Kruskal–Wallis ). Added validation yields no further gain, marking validators as reactive rather than proactive. Boundary experiments show tool quality binds only for in-distribution tasks with a canonical workflow; on out-of-distribution tasks hinging on a methodological decision, inter-strategy variance persists () and a runtime oracle revives the strategy advantage () — so the governing boundary is oracle IID/OOD, not bio-vs-general. We package these diagnoses into a reusable protocol whose transplant uplift self-targets lower-quality agents (, ).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.