CryptoDojo: A Framework for Safety and Security Evaluation of Crypto Agents
Abstract
Crypto agents translate user instructions into transfers, spending permissions, and commitments that can persist long after the decisions that created them. An individually authorized action can become unsafe when state changes or other actions interact with it, and evaluating actions in isolation misses this. We introduce CryptoDojo, a framework that injects controlled interventions during execution and reconstructs their financial effects to grade both policy compliance and task completion, while agents, attacks, defenses, and tasks vary. On CryptoDojo, we propose the Temporal-Compositional Attack Model (T-CAM), which targets the state, authority, ordering, and shared-resource dependencies among authorized actions, and instantiate it as T-CAM Bench. Across four agent configurations on 200 paired instances, safe completion drops from 53.5–90.5% on ordinary tasks to 0.5–14.0% under attack. Adversarial violation rates differ widely (3.5–28.0%), but safe completion does not. In 138 of 145 violations, the requested operation settles in full, and the violation lies in its timing or state, such as payment after a deadline or after a revocation. Configurations that violate less also settle fewer operations, so their safe completion is no higher. Evaluations of crypto agents should measure settlement and policy compliance together and track authority exercised after the agent stops.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.