acceptodds
Under review as a conference paper at ICLR 2027

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

Abstract

A multidisciplinary tumor board meeting serves as a critical decision point along the cancer patient journey, particularly representing one of the last hopes when standard options are exhausted. Agentic biomedical reasoning in large language models (LLMs) holds promise for timely multidisciplinary cancer decision-making. However, the development of these models is fundamentally constrained by the lack of evaluation using patient cases brought before real-world tumor boards, with discussion trajectories that integrate multimodal clinical observations and longitudinal patient timelines. We introduce OpenTumorBoard, a real-world benchmark at an unprecedented scale, with 611 patient cases and 19,157 turns of cancer discussions across ten specialist roles, transcribed from 12,534 minutes of video recordings on YouTube. The benchmark evaluates (1) LLMs joining a single Specialist Turn in the real discussion by responding to a clinically significant question; (2) Board Simulation of an entire back-and-forth discussion to reach a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Extensive evaluation of 14 general-purpose frontier models and medical LLMs identifies substantial limitations. The best-performing models reach only 3.43 out of 5 in clinical equivalence to actual specialist answers and 2.78 out of 5 in alignment with true tumor board conclusions. The best-performing models struggle both to answer naturally occurring questions posed by specialists and to orchestrate coherent multidisciplinary discussion trajectories. Supervised finetuning and reinforcement learning yield substantial improvements on a held-out test set, suggesting that OpenTumorBoard can serve as a training ground that effectively adapts LLMs with optimized clinical discussions. To confirm the quality of the benchmark, we recruit three M.D. experts and verify high information coverage and factuality of patient cases, as well as strong fidelity of the final consensus. We will release both OpenTumorBoard and its automated curation pipeline to enable scalable development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.