AlignDiag: Stage-Aware Multi-Source Diagnosis for Cross-Hardware Training Migration
Abstract
Large language model training increasingly spans accelerator generations and vendors as model development outpaces datacenter renewal. Common frameworks make such migration practical, yet backend-specific implementations can introduce legitimate numerical differences under finite-precision arithmetic. Meanwhile, similar losses may conceal internal faults whose effects propagate through subsequent computations. Reliable diagnosis must distinguish permissible drift from material errors and separate their sources from downstream effects. We introduce AlignDiag, the first staged multi-source diagnostic framework for training migration. It builds a contract-grounded evidence library that distinguishes noise, permitted drift, and controlled faults while preserving their provenance and lineage. Building on this structured evidence, a label-efficient procedure selects role-aware rules without using evaluated metrics to define labels. With these rules fixed, AlignDiag then detects and localizes material divergence automatically and uses targeted replay to verify whether identified candidates are genuine sources of observed deviations. Across our experiments, we organize 5,861 comparison pairs into 844 controlled groups covering expected variation and injected faults. Overall, AlignDiag achieves 98.61% accuracy distinguishing drift from faults. In prospective evaluations spanning training windows from 2 to 1,000 steps across GPU and NPU platforms, AlignDiag detects faults missed by loss-only checks, recovers injected source sets, and exposes sustained routing divergence in MoE models despite near-zero training loss. Multi-source localization reduces scans by 36.0% over iterative single-source strategies while preserving complete recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.