acceptodds
Under review as a conference paper at ICLR 2027

Fixed Activation Bounds Are Not Enough: A Distortion-Matched Transformer Study

Abstract

There are increasing societal concerns regarding safety around AI. Safe AI requires accurate predictions that withstand adversarial perturbations. Fixed bounds on internal activations constrain representation magnitudes, but whether they improve decision robustness remains uncertain. Here we show that fixed, training-calibrated hard and smooth caps preserve clean accuracy without establishing the prespecified practical robustness benefit in binary DistilRoBERTa classification on the Internet Movie Database benchmark. Across 55 matched training runs, we compare five transformer locations while matching mean absolute intervention on training-only activations; every nontrivial hard-cap distortion admits a unique positive smooth scale. All ten capped configurations satisfy corrected clean balanced-accuracy equivalence within one percentage point. At the primary embedding-space budget, none meets robustness-benefit criteria. Among 27 endpoint contrasts for caps with recorded activity, 15 establish corrected equivalence and 12 remain unresolved; the historically inactive smooth-logit comparator adds three equivalence decisions. At the less saturated budget, a post-hoc paired analysis likewise finds no corrected positive cap-control difference: hard query/key capping is 5.56 percentage points worse, seven contrasts remain unresolved, and two are equivalent. A contextual FreeLB-style baseline confirms that the evaluator distinguishes a more robust trained system. Our findings show a clear boundary between fixed activation control and robustness gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.