Omni Group Fusion: Groupoid-Compatible Correspondence and Equivariant Transport Across Multimodal Sensor Frames
Abstract
Omni and any-to-any models increasingly ingest text, images, video, audio, depth, and 3D into unified sequence backbones, achieving architectural interface uniformity through unconstrained attention. However, this convenience overlooks a fundamental physical reality: physical sensors and spatial language naturally operate in distinct, uncalibrated, and uncertain local reference frames. Treating directional and geometric quantities as flat, interchangeable scalars forces backbones to memorize pose-dependent statistical co-occurrences, precipitating sample inefficiency, cross-modal gradient starvation, and fragility under novel viewpoints or clock drift—while heuristic point-pose canonicalization breaks discontinuously under physical ambiguity. We introduce Omni Group Fusion (OGF), which grounds multimodal fusion on an observation groupoid respecting independent local coordinate charts. OGF operationalizes four transformation-consistent primitives: (i) routing tokens to shared world anchors via group-invariant descriptors, preventing correspondence drift under sensor rotations; (ii) maintaining multi-hypothesis spatial and temporal posteriors to represent frame ambiguity and clock drift without discontinuous collapse; (iii) equivariantly transporting typed representations ( vectors, tensors, coordinates) into a common world frame before pooling; and (iv) adjudicating consensus to attenuate conflicting sensors, updating anchors via Clebsch–Gordan couplings and reading back into frozen backbones through zero-initialized residuals. We theoretically establish global equivariance, local-gauge independence, stabilizer well-definedness, posterior composition consistency, and streaming stability. Across a controlled physical world, real continuous sensors, and five frozen generative backbones (3B–8B), factorial ablations isolate different claims: removing the group returns chance; typed transport, not orbit coverage, is what continuous sensors require; and, on Qwen2.5-Omni-7B, breaking sensor pairing removes routed synergy on four of the six audio–visual benches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.