Better Fixed-State Fidelity Need Not Improve Closed-Loop Performance in Quantized Neural Policies
Abstract
Does closer action agreement with a full-precision policy imply better closed-loop performance after quantization? We hold trained source policies fixed, vary calibration data or compiler settings, and execute compiled policies on Edge TPU and Hailo accelerators in simulated control tasks with paired initial conditions. In a complete, prespecified eight-rule Hailo FetchReach family, one artifact has a 16% smaller fixed-state mean action-error norm than another yet incurs a 34- fold larger estimated mean return loss relative to FP32. An early–middle contrast selected from that family and locked before replication favors early-state calibration in seven of eight independently trained FetchReach policies, and outcomeinformed audits retain large calibration effects under alternative compiler settings. Within the original family, a full-budget, AdaRound-enabled compiler intervention reduces action root-mean-square error (RMSE) on calibration-disjoint states for all eight rules, but mean return decreases for five; retrospective family-adjusted intervals place all five changes below zero but within the prespecified 0.5-return practical band. In a separate prospective benchmark of eight fresh FetchReach policies and 128 Hailo artifacts, a secondary analysis finds lower RMSE on a calibration-disjoint panel in 53 of 64 policy–rule comparisons, 22 of them with lower observed mean return; these are descriptive contrasts clustered within policies. Post-outcome episode resampling retains 16 of the 53 RMSE improvements with adjusted RMSE-change intervals below zero; the count varies across bootstrap seeds. Changing action coordinates or evaluation states changes which comparisons reverse, and reversals also occur when both executables are scored on the same visited states. We provide an execution-bound qualification protocol and provenance-recording harness that compare paired task outcomes against task-specific tolerances. The results distinguish fidelity-based candidate selection from task-tolerance approval and support evaluating the measured executable on its intended execution path rather than fixed-state fidelity alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.