Where Should Multimodal Test-Time Adaptation Capacity Live? A Controlled Study of Prompts and Residual Adapters
Abstract
We study where to place a fixed online adaptation budget in a frozen audio-visual transformer. Multimodal test-time adaptation must respond to audio-only, video-only, and paired shifts, yet prompt capacity and residual adapter capacity impose different update structures. We isolate this choice under a fixed BriMPR objective and a near-matched trainable-parameter budget, comparing prompt capacity with identity-initialized Prompt+Adapter capacity. The adapter is exactly identity at initialization, while a two-step integrity test verifies the intended gradient flow. Under a validated batch-32 severity-5 common-batch protocol (48 cases: 2 datasets 3 corrupted modalities 8 methods, three target-stream seeds), Prompt+Adapter attains the highest mean accuracy in all six dataset-modality settings, improving over prompt-only capacity by 0.04–2.41 percentage points and over the unadapted source model by 8.28 points on average. The ranking is stable across severities 1, 3, and 5, and placement and width pilots show that visual-side adapters match dual-side adapters while wider capacity helps both families. We deliberately scope our claims: we make no claims of universal superiority, causal mechanisms, or deployment efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.