NoRA: Learning Noise-Resilient Representation for Real-World Depth Perception
Abstract
Simulation enables scalable robot learning, but policies trained on pristine depth can fail under sensor-dependent noisy depth corruption in the real world. Restoring each input adds a denoising stage before control. Inspired by humans achieving robust perception through stable visual cues rather than reconstructing every pixel, we exploit a fundamental physical fact: sensing corruption changes measurements without changing scene geometry. We introduce **NoRA** (**No**ise-Resilient **R**epresentation **A**lignment), which learns stable geometric representations for control without external depth restoration. Its *Noise-Manifold Projection* combines teacher-guided spatial weighting and low-rank projection to suppress sensing-sensitive variation while retaining stable teacher content. NoRA distills clean geometric knowledge, aligns paired clean-reference and real-noisy depth at the policy input, and freezes the complete operator for simulation-only policy learning and subsequent zero-shot real-world transfer. NoRA provides eight CNN and Transformer backbones for different deployment budgets. Evaluations on depth perception, robot manipulation and multi-camera deployment show improved noisy-depth transfer. Under the evaluated sensing perturbations, selective alignment improves noisy geometry and real control over full-feature alignment, at a modest cost in clean geometry accuracy, revealing a practical geometry–robustness trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.