GazeWarp: Gaze-Guided Adaptive Warping for Bandwidth-Efficient Multimodal Interaction in XR
Abstract
Emerging wearable mixed reality (XR) systems increasingly adopt edge–cloud architectures, placing severe bandwidth constraints on the transmission of egocentric visual data to backend models. At the same time, natural XR interaction often involves ambiguous query expressions, with eye gaze serving as a practical cue to user intent. This intersection creates a critical challenge for gaze-guided multimodal interaction in XR: how to preserve the essential visual evidence needed to resolve a user's ambiguous query expression without exceeding a limited transmission budget. To study this problem, we introduce GazeRefer-GQA, a benchmark derived from GQA that combines user-view gaze cues, ambiguous query expressions, and structured reasoning annotations. We further propose GazeWarp, a gaze-guided visual enhancement framework that acts as a plug-and-play frontend before a frozen backend model. Using the user's gaze and question semantics, GazeWarp locates the intended target region and any structurally related regions, and drives a density-equalized warping process that spatially redistributes image resolution toward these regions before transmission. Across a range of transmission budgets and gaze input forms, GazeWarp generally outperforms resizing, cropping, and compression baselines under approximately matched payload budgets. Notably, despite using only a fraction of the bytes, it is comparable to the unconstrained full-resolution reference and numerically exceeds it in several settings. These gains hold across multiple backend vision-language models, and GazeWarp further generalizes strongly to additional gaze-guided datasets beyond our benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.