When Cross-Modal Evidence Helps: Modality-Local vs. Cross-Modal Representations for Reliability Inference in State Estimation
Abstract
Learning-based state estimation increasingly relies on multimodal sensing under changing sensing conditions, yet it remains unclear when reliability should be inferred from a sensor's own history and when cross-modal evidence is necessary. We introduce a Local–Cross framework that separates modality-local temporal representations from cross-modal relational representations while keeping the prediction target and model capacity matched. Using UAV optical flow, IMU, and GPS state estimation, we study two estimator-facing reliability quantities: covariance for stochastic uncertainty and consistency for compatibility with the nominal sensing model. We find that covariance prediction consistently favors modality-local representations, whereas consistency prediction consistently benefits from cross-modal representations, both when each degradation occurs alone and when the two occur simultaneously. When integrated into a classical ESKF, modality-local covariance prediction combined with cross-modal consistency prediction improves state estimation under degraded sensing while preserving nominal behavior. These results suggest that cross-modal evidence should be used selectively according to the reliability quantity being inferred rather than by default.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.