CALM-VR: Candidate-Specific Temporal Readout for Early VR Target Prediction
Abstract
Early target prediction in virtual reality (VR) ranks changing legal objects from causal multimodal histories. We ask whether candidate-specific temporal readout adds value when candidate information already enters both the final scorer and aligned history summaries. CALM-VR encodes gaze, head, and controller histories with modality-specific GRUs, then uses each candidate token to retrieve temporal evidence for masked ranking. A parameter-matched broadcast-query control shares one readout while retaining individual candidate tokens and aligned summaries; crossed controls vary summary alignment separately. An event-aligned, three-task collection provides mutually exclusive 97/21/30 fitting/selection, development/calibration, and frozen-protocol evaluation roles. In the 30-participant evaluation, 500-ms Top-1 was 51.20% versus 46.42% for designated late fusion (+4.78 pp, 95% CI [2.88, 6.68], Holm p < .001). The prespecified secondary query comparison estimated a +2.40 pp overall gain (95% CI [1.35, 3.45], Holm–3 p = .027). Equal task weighting gave a descriptive +1.40 pp point estimate; task-specific contrasts ranged from −3.56 pp in target acquisition to +6.15 pp in sequential assembly. Smaller positive aggregate gains were observed at 800 and 1000 ms. These results quantify the incremental value of candidate-specific temporal readout beyond candidate-conditioned scoring and aligned summaries, with observed variation across tasks in the evaluated VR setting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.