A³Flow: Attention Residuals, Adaptive Uncertainty Gating, and Motion Aggregation for Unsupervised Optical Flow
Abstract
Despite the dominance of RAFT-style recurrent correlation models, unsupervised optical flow methods with uncertainty estimation often inherits three restrictive design choices: fixed feed-forward residual encoders lacking data-dependent routing across scales, a hand-crafted stop-gradient gate that suppresses rather than correct updates when uncertainty is miscalibrated, and local correlation lookup that struggles to provide reliable correspondence cues in occluded regions. We present A³Flow, which augments an uncertainty-aware recurrent unsupervised framework with three pathways: (i) AttnResEncoder, which replaces fixed residual fusion between encoder blocks with learned softmax attention over pooled preceding-block representations, drawing inspiration from depth-wise attention in large language models to enable cross-scale routing in CNN encoders; (ii) ResidualFiLM(Residual Feature-wise Linear Modulation), which generalizes the fixed one-sided sigmoid gate to a two-sided bounded affine modulation driven by predicted uncertainty, enabling richer uncertainty-conditioned correction rather than one-sided suppression of flow updates; and (iii) CP-GMA, a Content-and-Position Global Motion Aggregation module integrated into the motion encoder, which propagates motion features from confident to occluded pixels, providing global motion context beyond local correlation lookup. Extensive experiments on MPI-Sintel and KITTI demonstrate that our method achieves new SOTA among unsupervised methods, with pronounced gains in occluded regions while preserving highly reliable uncertainty estimates. Code will be made available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.