Silent Collapse: Diagnosing and Repairing Low-Precision Robot Policy Training at the Action Level
Abstract
Robot learning is now trained largely with low precision, bf16 and fp8 alike. Low precision removes information the policy needs during training, and this can be silent: every monitor reads normal, yet the robot no longer produces the right action. We study this failure at the level of the executed action. It is not one failure mode but three, concentrated where the signal is weak along the activation gradient chain. The two most dangerous ones are new: mantissa starvation and conservative collapse. Both rewrite the action output without warning, with striking unevenness across tasks and seeds. We build a seconds-long probe locating the weak-signal points where failures start and present CURE (Cheapest sUfficient REpair), a routing protocol assigning each failure mode its cheapest sufficient fix: a free analytic bypass where the weak signal has a closed-form gradient, exact backward fidelity from the UPH (Unified Precision Health) layer where it does not, and LockedAdamW, an optimizer locking the per-step noise-displacement budget. We test the routing across policy families from a 0.8M ACT to 7B vision-language-action models under the twin-run protocol. On the ACT collapse cell, the routed scheme brings the collapsed posterior’s KL back into the fp32 band; on the 7B CogACT, LockedAdamW breaks below the fp32 self-retraining floor on both seeds; on real training hardware, one tested Transformer Engine fp8 configuration trains normally—consistent with the same mechanism—while unscaled variants fail in simulation. The probe answers in 3.4 seconds at 0.8M and in minutes at 7B; the repair starts free.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.