acceptodds
Under review as a conference paper at ICLR 2027

SHERLOC: Statistical Hypothesis Evaluation and Recovery for Learning Opponent Characteristics

Abstract

When used in automated decision-making systems, machine learning (ML) models are vulnerable to data-manipulation attacks. Some defense mechanisms (e.g., adversarial regularization) directly affect the ML models while others (e.g., anomaly detection) act within the broader system. However, many defenses operate under a predefined threat model that assumes characteristics of the adversary, including their knowledge, capabilities, and objectives, are known and fixed. In practice, these characteristics may be unknown or evolve over time, creating a mismatch between the assumptions of the defender and the true attacker. In this paper, we study a new perspective on adversarial defense: focusing on the attacker, rather than the attack. We present and demonstrate SHERLOC, a framework for adversarial learning forensics – namely, a detective that learns about the attacker from observed attack(s). We prove that, without additional knowledge, the attacker is non-identifiable — multiple potential attackers would perform the same observed attack. This reveals an important distinction between reconstructing an attacker's behavior and recovering the attacker that generated it. We demonstrate empirically that SHERLOC is able to reconstruct a hypothesis attacker that is consistent with the observed attacks against various defender models (i.e., generalized least-squares linear regressors, logistic regressors, and polynomial classifiers) across various security-critical domains. Finally, we show that these inferred attackers need not match the ground truth attacker exactly to have downstream utility. Defenses using the SHERLOC-inferred attackers achieve approximately of robustness improvement that would be achieved if the defender had access to the ground-truth attacker.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.