acceptodds
Under review as a conference paper at ICLR 2027

Towards Multi-View Sign Language Understanding: A Benchmark Dataset and Baseline

Abstract

Most existing Sign Language Understanding (SLU) methods are developed under fixed-view settings and remain vulnerable to viewpoint changes, while the scarcity of multi-view data hinders systematic research on viewpoint robustness. To address this gap, we introduce MVSign, a benchmark spanning three sign languages and seven viewpoints. MVSign combines original frontal recordings from existing datasets with additional views synthesized via D whole-body motion reconstruction and target-view rendering, while preserving the source splits and annotations. We further propose CanonSLU, a framework for multi-view SLU. This framework employs a canonical-view-guided strategy, using the frontal view as a semantic anchor. Specifically, CanonSLU incorporates Canonical-View Knowledge Transfer (CKT) to transfer canonical knowledge from the frontal view to non-frontal views, and Motion Relation Aggregation (MRA) to aggregate motion-related features across frames and strengthen temporal representations under viewpoint changes. Extensive experiments on MVSign demonstrate competitive performance with gains on non-frontal views, establishing CanonSLU as an effective baseline for multi-view SLU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.