BRACE: An R Benchmark for Rigorous Biomedical Data Science Code Search
Abstract
Efficient code retrieval is critical for biomedical data scientists, who navigate thousands of R packages on CRAN and Bioconductor and a growing body of open-source analysis pipelines. Existing code search benchmarks, however, focus on Python and rarely stress-test robustness beyond superficial lexical cues. To address this gap, we adapt an automated benchmark-construction pipeline to R and present BRACE (Biomedical R Anonymized Code rEtrieval), a benchmark built from real-world biomedical repositories and packages. BRACE contains 1,208 query-code pairs for evaluation and 5,310 pairs for training. Queries are LLM-generated natural-language descriptions validated through scoring by domain experts and hypothesis testing. The pipeline first ensures that every snippet byte-compiles and resolves all of its dependencies inside a pinned R/Bioconductor environment. It then categorizes snippets by dependency complexity, distinguishing functions that use only base R, functions that rely on custom S3/S4/R6 or Bioconductor container classes, and functions that invoke user-defined helpers. BRACE further stress-tests retrieval robustness through identifier anonymization and through representation shift, in which functions are lowered to R bytecode or serialized as S-expression parse trees. Under these conditions, our evaluation of six retrieval models reveals consistent drops when identifiers are anonymized, NDCG falls by 3-25 points on base-R functions and by up to 56 points on container-class functions, and far larger drops on lowered representations, where even the strongest model ranks the correct function first for fewer than a quarter of bytecode queries. The results indicate that current models still rely on lexical features rather than code semantics, and that in R this reliance extends to string-typed data accessors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.