acceptodds
Under review as a conference paper at ICLR 2027

INVARIANCE DOES NOT GUARANTEE REPAIR IN FROZEN-FEATURE DETECTORS

Abstract

Removing a nuisance subspace from frozen features makes a detector exactly invariant to the removed directions. Whether that repairs the group error that motivated the removal is a different question, which we study in authenticity detectors on frozen speech and image encoders. Deleting a rank-4 emotion subspace estimated by the same procedure on acted corpora raises false rejection of surprised genuine speech from 17.3% to 33.9% on a w2v-BERT 2.0 head and lowers it on a WavLM head. What reverses the outcome on that w2v-BERT head is where the eraser is estimated. With four or five labels and matched rank, every tested eraser estimated on acted corpora raises the error, the same constructions estimated on genuine target-domain speech lower it at a cost in neutral rejection, and that speech helps more still when it trains the detector instead. The removed rank, the residual decodability of the nuisance and its mean response are three diagnostics an erasure audit can report, and none of them decides between the two outcomes. For a fixed linear head, the removed part of that response gives the direction in which the response moves but not reliably that of the error. Controlled image experiments show the limit of the mean response as well. With the audited geometry fixed, a second training association changes the group error several-fold while the mean nuisance response remains similar. Repair therefore has to be judged on target-domain decisions, for the genuine groups and generated sources at stake, not on the erasure itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.