acceptodds
Under review as a conference paper at ICLR 2027

Where Credit Lands: Verifiable Step Credit for Bounded-Memory EHR Agents

Abstract

Reinforcement learning with verifiable rewards (RLVR) has enabled large language model (LLM) agents to solve clinical tasks over longitudinal electronic health records (EHRs), where an agent takes a sequence of retrieval and memory decisions over evidence scattered across years of visits. However, the usual optimizer, group relative policy optimization (GRPO), reduces each rollout to one score and gives every action the same advantage, so it cannot tell which action helped or hurt. Crediting each action needs its own score and a comparison with its alternatives, neither of which is readily available for current EHR agents. We propose Verifier-Anchored Process Advantage (VAPA), a critic-free training method that pairs checks of individual actions with comparisons of alternative decisions, needs no learned reward model, and uses neither verifiers nor replay at test time. First, deterministic checks score retrieval, memory, and stopping actions from the patient record, completed rollout, and task specification. Second, forked replay rewinds to the state before a selected action and samples alternative continuations, with weaker fallback groups where no fork reaches. Finally, a two-level advantage assigns local credit from these comparisons while preserving final-answer credit for every original action. Across MIMIC-IV calculation and retrieval tasks and the held-out MedAgentBench and EHRSHOT suites, VAPA has the highest mean among matched agent, memory, and RL baselines, with no further training for transfer. At the same sampled-token budget, replaying alternative decisions outperforms additional independent rollouts, and VAPA's lead over trajectory-level GRPO grows with decision-chain length and with a tighter memory bound.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.