BAViR: Aligning Urban Visual Representations with Revealed Mobility Behavior
Abstract
Urban environment representations should describe how places are used, not only their visible structure. Mobility records reveal this use but need not be available wherever a representation is deployed. We use mobility behavior as training supervision to learn an environmental representation that can be applied where mobility observation are unavailable. We introduce BAViR (Behavior-Aligned Visual Representations), which fuses satellite and street imagery with points of interest and road structure, and a contrastive objective aligns its output with visitation, dwell, activity, and temporal summaries. Since mobility behavior is used for supervision only during training, the learned encoder can represent new locations from environmental observations alone. We evaluate spatial generalization and reuse for targets excluded from encoder training, using two five-fold spatial protocols and three seeds in metropolitan Orlando. On geographically withheld regions, BAViR improves visit-count from 0.488 to 0.523 over matched masked-stream self-supervision. Moreover, the learned representation also transfers to transaction spend, a target excluded from encoder training, reaching of 0.355 compared with 0.314 for self-supervision and 0.339 for behavior regression under the same frozen-feature readout. With only 10% of downstream spend labels, BAViR exceeds raw-feature Ridge trained with all labels. Together, these results support the value of behavior supervision for learning reusable functional representations, rather than a universal advantage of contrastive learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.