CLINCH: Linking Evidence, Decisions, and Execution in Long-Horizon Clinical Tasks
Abstract
Clinical agents can make locally appropriate decisions yet still fail to complete long-horizon tasks in electronic health record (EHR) systems. Completing such tasks requires gathering patient evidence dispersed across records, revising decisions as that evidence changes, and distinguishing intended actions from operations actually completed, so that the final documentation reflects what was done. We introduce CLINCH, a clinical agent harness that maintains a revisable workspace linking task requirements, patient evidence, clinical decisions, and execution status. CLINCH preserves and recovers decision-relevant context, provides patient-specific medical knowledge at decision points, and checks task completion and documentation against runtime-recorded actions and results. Identified gaps return to the agent for further retrieval, execution, or revision. We evaluate CLINCH with DeepSeek-V4-Pro, GPT-5.6-Luna, and GLM-5.2 on PhysicianBench and MedAgentBench. Relative to the corresponding baseline agents, CLINCH improves task success by 13.0, 4.7, and 21.0 percentage points on PhysicianBench and by 11.1, 12.2, and 19.4 points on MedAgentBench, respectively. On PhysicianBench, it also reduces failed checkpoints in retrieval, clinical reasoning, action execution, and documentation for all three models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.