MASK-GATE: Measuring and Mitigating Cross-User Leakage in Evolving Agent Skills
Abstract
Agent skills increasingly distill execution experience into reusable instructions that improve future agent behavior. Reusing these evolved skills across users introduces a new risk: information from one user’s trajectories may persist in a shared Skill and influence another user’s decisions even when model weights remain fixed. We treat the trajectory-to-shared-Skill interface as a controlled information-release boundary that preserves task-general evidence while excluding user-scoped signals. We introduce MASK-GATE, an evolution-time framework that enforces this boundary through typed semantic masks, sanitized Evidence records, Evidence auditing, and fail-closed privacy, integrity, and capability checks. These controls operate offline, add no model calls at inference time, and retain the Original Skill when an update is rejected. On CrossUserSkillLeak, a controlled benchmark for cross-user leakage in location-based tool use, we compare the no-evolution Original Skill, SkillOpt, SkillClaw, and MASK-GATE under the same registered evolution groups, frozen base model, tool snapshot, probe inputs, and evaluation protocol. The 100-probe test set is held out throughout method development and is opened only after all artifacts, release decisions, and evaluation rules are frozen. Across 165 registered evolution groups and 100 held-out probes, MASK-GATE achieves the highest observed grounded task correctness, reaching 74.0%, compared with 68.9% for SkillOpt, 65.7% for SkillClaw, and 51.0% for the no-evolution Original Skill. It also attains the lowest behavioral contamination rate (BCR), at 0.9%, compared with 2.8% for SkillOpt and 5.7% for SkillClaw. Our method therefore attains the lowest observed leakage and highest observed correctness among the evaluated methods, with gains concentrated in capability-generalization probes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.