acceptodds
Under review as a conference paper at ICLR 2027

Reject Classes and Anchored Energy: Bypassing the Clean-Robust Trade-off in Adversarial OOD Detection

Abstract

Neural networks are often overconfident on out-of-distribution (OOD) inputs, a vulnerability exacerbated by adversarial perturbations designed to evade detection. We reveal that In-Distribution (ID) adversarial training is indispensable for genuine OOD robustness—without it, even the strongest OOD-centric defenses collapse under principled attacks. However, it induces a severe clean-robust trade-off: ID adversarial training suppresses the maximum softmax probability (MSP) of clean samples, crippling the separation between ID and OOD data. We establish two structural routes to mitigate this trade-off: (1) replacing MSP with the LogSumExp (LSE) energy score and supplying its missing anchor via energy-margin regularization (RCE+EMR); and (2) GRoOD, which unlocks an auxiliary reject class whose probability stays insulated from ID confidence degradation. We then uncover a fundamental bottleneck when attempting to transfer EMR across the two routes: the variance-free vulnerability—EMR fails on GRoOD because its OOD objective supervises only the reject class, leaving the internal logit variance unconstrained and exposing the LSE operator's extreme-value sensitivity to single-dimension adversarial breakthroughs. Introducing OOD label smoothing to architecturally lock this missing variance unleashes EMR's full potential on GRoOD, yielding GRoOD+OOD-LS+EMR that attains the best robustness in reject-dominated regimes while providing both detection signals in a single model. Together, the two structural repairs cover both signal regimes, with RCE+EMR leading where LSE detection dominates.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.