OrgMem: Benchmarking Self-Evolving Agents for Organization-Specific Evolving Audit Environments
Abstract
Large language model (LLM) agents increasingly rely on external memory to reuse experience across long-horizon interactions, yet existing evaluations provide limited insight into whether historical experience remains useful when the underlying decision environment evolves. This challenge is particularly important in organizational auditing, where procurement policies, market conditions, vendor behavior, and fraud-risk patterns may change over time, causing superficially similar cases to require different judgments. We introduce AuditEvo, a synthetic benchmark for evaluating self-evolving agents under organization-specific and implicitly evolving procurement environments. AuditEvo contains 40,000 evaluated cases organized into 100 chronological organizational chains with evolving policy, market, vendor, and risk conditions. At each step, an agent observes the current case and chronologically available history, predicts an audit outcome of approval, rejection, or investigation, and immediately receives verified feedback that may be reused in subsequent decisions. A deterministic audit oracle provides reproducible labels while keeping latent environment states and oracle rationales hidden from evaluated agents. Experiments across diverse memory and reflection frameworks and multiple LLM backbones show that current systems struggle particularly with risk-regime cases and that retrieving historically related cases alone does not ensure reliable decisions when their applicability changes. Motivated by this limitation, we introduce Environment-Aware JitRL, which maintains an evolving playbook that uses feedback to adapt state similarity and thereby condition historical experience retrieval for action-value estimation. Across evaluated backbones, the method improves average audit accuracy over JitRL by approximately 2–7 percentage points. AuditEvo provides a controlled testbed for studying memory reuse, adaptation, and negative transfer under persistent multi-environment evolution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.