acceptodds
Under review as a conference paper at ICLR 2027

How Agents Reach the Record Matters: A Parity-Controlled Study of Access Modality for Clinical EHR Agents

Abstract

How access modality affects the measured performance of agents that act on electronic health records (EHRs) remains poorly characterized. We evaluated five frontier models from OpenAI, Google and Anthropic on ten EHR tasks, from chart retrieval to consults with orders and signed notes, under three access modalities: Pixel (screenshots, mouse and keyboard), FHIR (patient-scoped FHIR search and read) and Packet+Action (the full record in the prompt, with FHIR's actions). Tasks, synthetic records, actions and grading were held constant under a protocol frozen before any model call. Physicians wrote or reviewed each task's clinical content. Grading combines deterministic checks on the final chart and recorded actions with a blinded judge for consult notes. Across 1,500 trials (10 per cell), pass@1 was 32.4% under Pixel, 54.0% under FHIR and 60.0% under Packet+Action. Exact task-level sign-flip tests place Pixel 21.6 points below FHIR (p = 0.031) and 27.6 below Packet+Action (p = 0.008); the structured modalities do not differ reliably (p = 0.109). Every model performed worst under Pixel, between-model differences were largest under Pixel (6.0% to 48.0% pass@1), and modality explains about as much variance in pass rates as model choice. Most of the Pixel deficit arises in reaching evidence and in finishing and persisting writes; for one model it may largely reflect the screen driver. Auditing our harness exposed parity defects, and every affected cell was rerun under a frozen plan. These results establish access modality as a critical dimension of EHR-agent evaluation, and we state the parity requirements such comparisons need.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.