From Safety Notices to Monitoring Policies: Evidence-Grounded LLM Agents for Statistical Process Control
Abstract
Safety notices require hospitals to map natural language onto local exposure cohorts before statistical monitoring can begin. We formulate this step as a probability distribution over executable cohort predicates, linked to monitoring actions by an equal- expected-loss solver. An LLM uses six calls to five read-only tools to acquire cell-suppressed aggregate evidence and update its cohort distribution. Deterministic code handles calibration and action selection, while patient rows, outcomes, and monitoring losses remain outside the model's interface. We evaluate 56 synthetic notices for 14 cohorts held out from semantic-method development, using cohort memberships derived from real MIMIC-IV exposures and two prespecified process-shift templates. The Agent reduces normalized regret by 38.6% relative to the same LLM without tools and by 64.9% relative to a fixed-query non-LLM pipeline with the same six-call budget. Mean detection delay falls to 53.92 patient arrivals from 58.24 and 67.14, respectively. Fifteen of sixteen preregistered conditions were met; the largest target's 27.86% share of net gain against fixed-query evidence exceeded the 25% concentration ceiling, leaving the conjunctive confirmation criterion unmet. These results quantify how model-directed aggregate evidence use improves cohort-conditioned monitoring decisions within a fixed executable catalog.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.