TRACE: FROM VERIFIED OUTCOMES TO REUSABLE PROBING STRATEGIES FOR MEMORY-EXPOSURE ASSESSMENT IN CLINICAL LLM AGENTS
Abstract
In clinical consultations, LLM-based clinical agents may use patient-specific records while also retaining information from earlier interactions, including prior dialogue, retrieved electronic health record context, and tool-use traces. This persistence introduces a privacy risk: retained patient information may later reappear in a response even when it is absent from the patient's current query. We refer to this response-level resurfacing of retained information as memory exposure. In this paper, we study whether an external assessor can acquire reusable probing strategies for systematically assessing such exposure without modifying the target clinical agent. The assessor probes the target through ordinary interactions. A probing strategy is a reusable pattern for determining the assessor's next query or action. Thus, we propose TRACE, a verifier-guided framework for the bounded evolution of reusable probing strategies. TRACE starts from a fixed seed repertoire and executes the available probes against the target agent. It uses a deterministic verifier to compare each response with protected information defined for the trial. Exposure is counted only when the matched information is absent from the visible query. The resulting exposure and non-exposure outcomes are stored as structured records of the probe, target response, and verification result. These records guide the acquisition of new probing strategies. A response-conditioned router then selects among the retained strategies based on the target's observed response. After the acquisition phase, both the learned repertoire and router are frozen. During downstream evaluation, no new probing strategies are generated and the router is not updated, so the measured performance reflects reuse of acquired probing knowledge rather than continued test-time adaptation. Across five independently seeded acquisition runs, TRACE increases mean Verified Success Rate (VSR) from 62.0% with the fixed seed repertoire to 93.6% on the held-out evaluation split. With acquisition disabled, the same frozen repertoire and router achieve a mean VSR of 93.3% across DeepSeek-V3.2, GPT-4o, and LLaVA-v1.5-7B. Additional analyses examine sensitivity to routing, context construction, dataset shift, and output-side controls. These results support verifier-guided acquisition and frozen reuse of probing strategies for memory-exposure assessment in clinical LLM agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.