Cross-Scale Model Resilience for LLM Inference Under Hardware Faults
Abstract
Modern large-language models (LLMs) deployed on computing systems at scale are increasingly exposed to transient faults. These faults can silently corrupt persistent model state (i.e., parameters), affecting subsequent inference and potentially causing consequential decisions and actions. Existing resilience mechanisms detect and correct corruptions in bits, activations, and parameters. However, bounding the numerical errors from such corruptions does not guarantee semantically correct outputs. In this work, we present cross-scale model resilience (CSMR), a mechanism that detects and mitigates fault-induced errors at the output level. CSMR pairs a target model with a lightweight auxiliary model and measures their token- and sequence-level semantic discrepancy, triggering correction only when the discrepancy exceeds a calibrated threshold. To minimize overhead, we integrate CSMR into speculative decoding, using the auxiliary as the drafter and embedding discrepancy checking into the acceptance stage. Upon detection, CSMR selectively reloads corrupted parameters. Across multiple LLMs and benchmarks under diverse fault models, CSMR restores accuracy from 0% to 95.6–99.6%, with less than 0.4% latency overhead in fault-free executions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.