acceptodds
Under review as a conference paper at ICLR 2027

Beyond Average Success: Outcome Reversals and Delayed Divergence in Quantized Vision-and-Language Navigation

Abstract

Quantization is widely used to compress models, reduce memory requirements, and potentially accelerate inference. When average performance is preserved, a natural expectation is that quantization introduces only local numerical perturbations while leaving model outputs and most individual outcomes largely unchanged. However, our experiments in streaming vision-and-language navigation (VLN) reveal substantial behavioral differences behind similar average scores. Changed trajectories can also alter subsequent observations and histories. Using instruction-paired R2R-CE episodes, we compare four-bit weight quantization of StreamVLN against its unquantized bfloat16 (BF16) model as the full-precision baseline. Success rate decreases by only 0.11 percentage points, yet 18.5% of outcomes reverse and 87.0% of action sequences change, as gains nearly offset losses. To examine this discrepancy, fixed-state probes separate numerical perturbations from immediate decisions and distinguish the effects of current computation and historical content. Further, controlled continuations show that history perturbations can induce later divergence even when the initial action queue is held fixed and subsequent computation returns to BF16. Together, these findings show that similar average performance does not imply behavioral fidelity, and immediate action agreement does not guarantee subsequent agreement in the tested states. Behavioral differences need not reduce task success. Quantized VLN deployment therefore calls for evaluating retained task outcomes and subsequent behavior alongside average quality, realized memory savings, and end-to-end latency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.