acceptodds
Under review as a conference paper at ICLR 2027

MedBenchAgent: An Agentic Framework for Automated Medical Benchmark Construction

Abstract

The growing use of large language models (LLMs) in the medical domain drives demand for benchmarks to evaluate their capabilities. Constructing these benchmarks manually is labor-intensive, motivating the use of LLM-based agents to automate the process. However, general-purpose agents often lack sufficient medical expertise for benchmark construction, while existing medical benchmark agents typically target specific tasks or modalities and struggle to consistently produce high-quality benchmarks, limiting both generalizability and reliability. To address these limitations, we introduce MedBenchAgent, an agentic framework for reliable medical benchmark construction across multiple tasks and modalities. MedBenchAgent follows a two-stage workflow supported by an extensible medical tool library. The benchmark design stage decomposes the user request into subtasks, integrates subtask responses into a benchmark specification, and uses this specification to generate template items. To improve design reliability, a field-level quality model guides specification refinement. The data construction stage retrieves relevant medical samples for each template item, performs fine-grained annotation on the retrieved samples, and instantiates final benchmark items from the templates and annotated data. To improve annotation reliability, an evidence-guided Thought-Tool-Observation loop verifies and refines annotations. The medical tool library combines diverse medical resources and specialized tools with dynamic tool expansion to support both stages and improve generalizability. Experiments on five benchmark construction tasks spanning CT, pathology, medical video, electrocardiography (ECG), and electronic health records (EHR) show that MedBenchAgent outperforms existing methods overall across LLM jury evaluation, downstream fine-tuning, and human evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.