acceptodds
Under review as a conference paper at ICLR 2027

CRAFT: Coupled Rotation-Aware Action-Field Caching and Two-Bit Quantization for Vision-Language-Action Models

Abstract

Diffusion and flow-matching vision-language-action (VLA) policies incur substantial weight traffic through repeated action-head evaluations. Prior VLA quantization typically retains head weights at W4 or higher, while scalar W2 can severely degrade task success. Adaptive caching introduces per-step decision overhead that can offset savings on short trajectories (), despite often selecting identical evaluation steps. We introduce Coupled Rotation-Aware Action-Field Caching and Two-Bit Quantization (CRAFT), a training-free framework that jointly reduces action-head precision and evaluation frequency. Motivated by scalar-W2 error concentration in a few high-activation-energy directions, CRAFT quantizes the head to W2A4 using lattice codewords. It selects modules for rotation before deployment using the original weights' tail statistics and reuses the latest predicted field at fixed steps, skipping head evaluation while continuing solver-state updates. With a W4 backbone, W2A4 head, and caching, CRAFT achieves LIBERO-Object success rates of 98.4% and 97.8% on GR00T N1.5 and (FP16: 96.4% and 98.0%). Batch-1 end-to-end speedups on RTX 3090 reach and , respectively, versus and for the fastest prior methods. On GR00T N1.5, nominal weight traffic decreases by 79.6% relative to FP16 and by a factor of 2.1-3.2 relative to prior methods. Head-only W2A4 quantization achieves 80.45% average success across eight models, exceeding the strongest matched-precision baseline by 15.3 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.