SonoArena: Formulating and Benchmarking Agentic Decision-Making in Ultrasound Examination
Abstract
Ultrasound examination is a sequential decision-making process: a sonographer interprets the evidence acquired so far, decides where to scan next, and stops once the evidence is sufficient. Previous work on ultrasound AI has largely overlooked this decision loop. In this work, we formulate ultrasound examination as a partially observable Markov decision process (POMDP), in which diagnostic evidence is revealed only when the corresponding acquisition action is taken. Building on this formulation, we introduce SonoArena, an interactive benchmark in which ultrasound agents sequentially acquire diagnostic evidence and decide what to report. We theoretically characterize the operating regimes of different policy types, which are then supported by the empirical results. We benchmark 24 large language model (LLM)-based policy configurations, spanning general-purpose and task-trained policies. The strongest task-trained policy outperforms the evaluated baselines overall but remains below the reference, leaving room for more targeted training methods. These results demonstrate the necessity and feasibility of evaluating and training ultrasound agents not only by how well they interpret acquired observations, but also by whether they make the right acquisition decisions to support their conclusions. Code is provided in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.