acceptodds
Under review as a conference paper at ICLR 2027

TwinCheck: Comparative Safety Verification for Self-Evolving Agents

Abstract

Self-evolving agents improve by retaining feedback as reusable memories or skills, but every mistaken approval can shape many future behaviors. We identify a blind spot in this process: two updates can be identical with respect to all information used by the admission verifier yet produce opposite safety outcomes, so recalibra- tion cannot distinguish them. We therefore introduce TwinCheck, a verification framework that filters harmful updates before they can contaminate a self-evolving agent’s persistent state. Specifically, TwinCheck constructs an alternative that matches every signal used by the original verifier, without access to safety labels. It then replays both updates from the same state, compares the resulting behaviors in both presentation orders, and averages the two scores before deciding whether to retain the candidate. We evaluate TwinCheck on two paired-update evaluation sets built from AI Safety Gridworlds and τ 2-Bench, covering executable programs and natural-language memories. Across eight open-weight and proprietary-model verifiers, TwinCheck raises mean harmful-update detection AUC from 0.748 to 0.945 on controlled traces and from 0.695 to 0.835 on tool-use conversations, outperforming VaG and AgentDoG 1.5. In a ten-generation self-update loop, it reduces trajectories experiencing degradation from 30.39% to 0.43% and leaves no final degradation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.