QAttrib-MoE: Runtime-Controlled Attribution of Cross-Precision Trajectory Divergence in Mixture-of-Experts Inference
Abstract
Low-precision mixture-of-experts inference is usually compared by endpoint accuracy, yet similar accuracy can conceal substantial changes in the autoregressive trajectory, while runtime variability can confound attribution to numerical precision. We introduce QAttrib-MoE, a pre-frozen, runtime-controlled, noise-floor-aware protocol for attributing cross-precision trajectory divergence. Same-path fresh-server variability defines an empirical control floor, cross-path divergence is the observed contrast, and excess divergence measures the contrast remaining above the two path-specific floors. The protocol freezes the evaluation cohort and decision rule before generation, counterbalances execution order, and reconstructs multiple decoding budgets by censoring a single stored trajectory rather than regenerating requests. We instantiate QAttrib-MoE on a shared GLM-5.2 checkpoint executed through FP8 and online-NVFP4 routed-expert paths. At the pre-frozen 256-token prefix, both paths exhibit 0.00%/0.00% within-path mismatch on 64 controls, while cross-path mismatch and excess divergence reach 98.44% (95% CI: 95.31-100.00%); the effect is positive on both GSM8K and MMLU-Pro and grows from 50.00% at 32 tokens to 98.44% at 256 tokens. In contrast, strict 1024-token accuracy is 64.45% versus 66.41% (difference 1.95 pp, 95% CI: -1.95 to 5.86 pp; McNemar p = 0.442). A gated teacher-forced diagnostic preserves token-ID alignment yet finds top-1 prediction flips in all 48 anchored inputs, with median first flip position 2. The results show that individually reproducible precision paths can be strongly non-equivalent at the reasoning-trajectory level without a statistically significant endpoint-accuracy shift in the evaluated setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.