acceptodds
Under review as a conference paper at ICLR 2027

TriViS: A Large-Scale Multi-View Benchmark for Vietnamese Sign Language Recognition

Abstract

This paper presents TriViS, the first large-scale multi-view Vietnamese dataset for Continuous Sign Language Recognition and Translation, designed as a controlled benchmark for studying dataset design in underrepresented sign languages. TriViS contains 12,000 sentences and 84,453 synchronized videos captured from three viewpoints (frontal, left, and right) across both controlled studio recordings and outdoor environments. The synchronized multi view setup enables systematic analysis of viewpoint sensitivity, view complementarity, and other data collection factors that are difficult to isolate in uncontrolled web-scale corpora. Along with the dataset, we introduce a simple multi-view fusion framework that integrates synchronized visual streams into a unified representation for sentence-level recognition and translation. Using state-of-the-art CSLR and CSLT baselines, we conduct extensive experiments to evaluate the contribution of individual viewpoints, multi-view fusion strategies, sentence-independent generalization, and the effect of LLM scale. Our results show that multi-view observations consistently outperform single-view baselines and provide new empirical insights into how viewpoint configuration and controlled data collection affect continuous sign language understanding. Our code and dataset are publicly available on https://github.com/Anonymous-Submission-abcxyz2501/TriVIS_Submission_ICLR2027.gitGitHub and https://huggingface.co/datasets/Abcxyz2501/Trivis_Representative_SamplesHugging Face.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.