acceptodds
Under review as a conference paper at ICLR 2027

Fault Internalization: When Training Optimization Absorbs Injected Hardware Faults

Abstract

Outsourced AI model training and fine-tuning increasingly rely on third-party platforms, where trusted execution environments (TEEs) are widely adopted to isolate sensitive model weights and training data from the untrusted privileged host. However, hardware fault injection through bit-flip attacks (BFAs) remains capable of perturbing protected computation. Existing BFAs either require white-box access to model weights or destructively corrupt low-level execution, making them unsuitable for practical training-stage attacks. We uncover fault internalization, a training-specific phenomenon in which optimization adapts to altered computation by encoding compensation into the learned weights. An untrusted host can exploit this behavior to keep training and validation normal when performing BFAs, thereby concealing the ongoing attack. Once the model is transferred to a clean platform, the original computation is restored while the learned compensation remains encoded in the weights. The compensation then no longer offsets the altered computation and instead drives the model toward the adversarial behavior. We characterize the semantic- and implementation-level conditions that enable fault internalization and develop a systematic procedure for mapping suitable computation alterations to target bits in underlying computation libraries. We evaluate fault internalization on discriminative and generative models under training-from-scratch, full-weight fine-tuning, and LoRA fine-tuning. Across these settings, attacked execution remains close to benign behavior, whereas restoring the original computation on a benign platform exposes diverse model-level failures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.