Characterising Input Loss Curvature Under Backdoored Networks
Abstract
Backdoor attacks, introduced through data poisoning or architectural manipulation, induce attacker-chosen behaviour on inputs carrying a trigger while leaving clean model performance essentially unchanged. The eigenvalues of a model's input loss Hessian, , measure the local curvature of the loss surface around an input and may therefore contain information about triggers in backdoored networks. We study this through the exact decomposition , where is the generalised Gauss-Newton term and is a model-curvature correction. For cross-entropy classification, we derive bounds showing that output-space curvature collapses as prediction confidence approaches unity. In locally affine networks, almost everywhere, so input curvature also collapses if the model Jacobian remains bounded. We further show that locally affine additive backdoors only change , whereas multiplicative composition introduces an indefinite residual term of rank at most two. We validate these results experimentally using ResNet-18 image classification models with ReLU and SiLU activations. For additive backdoors, we observe eigenvalue collapse under ReLU activations that does not persist when SiLU is used, and multiplicative backdoors produce a large model-curvature correction. Overall, these results show that input curvature depends jointly on backdoor construction, model confidence, model architecture, and loss function.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.