acceptodds
Under review as a conference paper at ICLR 2027

What Can WiFi Say? Language-Aligned Representation Learning for WiFi Signals

Abstract

Understanding everyday activity could support smarter homes, personal assistants, and remote care. Camera-based monitoring provides rich observations but raises privacy concerns, while WiFi offers a privacy-aware alternative. We therefore ask: What can WiFi say in natural language? With language, WiFi could describe movement, retrieve related activities, and answer questions about what happened. We introduce SignalLang, a unified framework that aligns WiFi channel state information (CSI) with motion-focused language. Its captions retain posture, body-part involvement, motion attributes, and temporal structure while excluding unsupported visual details. A transformer over subcarrier-time tokens preserves the link, frequency, and temporal structure of CSI, while contrastive, matching, generative, and auxiliary activity objectives learn a shared WiFi-text representation. Class-aware negative masking prevents recordings of the same activity from being treated as negatives. A single backbone supports caption generation, bidirectional retrieval, and WiFi question answering. On XRF55 activities represented during training, SignalLang improves over an adapted WiFi2Cap baseline by 19.9 BLEU-4 points, 51.8 percentage points in WiFi-to-text retrieval R@1, and 25.3 percentage points in WiFi question answering accuracy. It also shows measurable transfer to held-out activity classes in retrieval and question answering.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.