Deep Learning-Based Sign Language Translation: A Survey and Taxonomy
Sign Language Translation (SLT) is crucial for communication between Deaf or Hard-of-Hearing communities and society. Advances in deep learning have enabled promising sign-to-text systems, yet progress remains constrained by limited data, complex spatiotemporal dynamics, and non-manual cues. This survey reviews deep learning-based SLT methods across gloss-based, gloss-free, and weakly gloss-free paradigms, summarizing reported performance on RWTH-PHOENIX-Weather 2014T (PHOENIX-2014T), CSL-Daily, and How2Sign. Gloss-free approaches are increasingly explored as annotation-efficient alternatives, leveraging contrastive learning, self-supervised alignment, and multimodal fusion to reduce reliance on costly manual gloss annotations. Remaining challenges include dataset diversity, temporal modeling, evaluation reliability, real-world robustness, and cross-linguistic generalization.
Authors
- Ruili Wang (ORCID: https://orcid.org/0000-0003-2899-9816)
- Desen Yuan
Institutions
- Massey University (NZ)
Publication Details
- Journal
- ACM Computing Surveys
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1145/3844609
- Primary Topic
- Hand Gesture Recognition Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00