ZORA: Zeroth-Order Reliability-Calibrated Adaptation with Shift-Sensitive Updates for Multimodal Test-Time Shifts
Abstract
Audio-visual continual test-time adaptation (AV-CTTA) enables pretrained multimodal models to adapt online to non-stationary target streams. However, most existing AV-CTTA methods rely on backpropagation-based updates, limiting their deployment on edge platforms with quantized backbones, inference-only engines, or strict memory and latency constraints. Zeroth-order optimization (ZOO) provides a forward-only alternative, but its direct application to AV-CTTA faces two challenges: zeroth-order queries must explore a high-dimensional multimodal parameter space without shift awareness, and entropy minimization may reinforce incorrect predictions that exhibit high confidence and low entropy under online multimodal shifts. To address these limitations, we introduce Zeroth-Order Reliability-Calibrated Adaptation (ZORA), a forward-only AV-CTTA framework that determines both where ZOO should update and which learning signals it should trust. ZORA incorporates Shift-Sensitive Layer Selection (SSLS) to restrict zeroth-order updates to shift-sensitive normalization layers in the audio and visual branches. It further employs Bias-Aware Reliability Calibration (BARC) to assign signed entropy weights based on prediction bias and confidence, while aligning target-window feature statistics with cached source-reference statistics. Experiments on Kinetics50-C and VGGSound-C demonstrate that ZORA consistently outperforms representative CTTA, multimodal adaptation, and forward-only baselines across multiple continual audio-visual shift protocols. Experiments with quantized CAV-MAE models and efficiency analyses further support resource-constrained forward-only deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.