Open-Domain Sign Language Translation Learned from Online Video
Bowen ShiDiane BrentariGregory ShakhnarovichKaren Livescu
Presents OpenASL, the largest open-domain American Sign Language dataset from online video, alongside pre-training and multi-feature fusion techniques that substantially improve gloss-free translation in unconstrained environments.
Automatic sign language translation is critical for improving artificial intelligence accessibility for more than 430 million deaf and hard-of-hearing individuals worldwide. However, existing research has largely relied on small datasets recorded in controlled studio settings or narrow domains, such as weather forecasts. Furthermore, prevailing translation models heavily depend on intermediate gloss annotations—transliterations that directly transcribe sign language—which are costly, difficult to scale, and rarely available in real-world scenarios.
The article introduces OpenASL, a large-scale American Sign Language (ASL) dataset, and evaluates a gloss-free translation framework designed to handle continuous signing directly from real-world video into written English sentences.
To build the dataset, the authors collected 288 hours of online videos from YouTube news channels and community video blogs, covering over 200 signers across 98,417 video-sentence translation pairs. Because the source videos were self-generated with aligned English captions rather than interpreted broadcasts, sentence-level boundaries showed high temporal alignment without requiring full manual re-captioning. For the translation architecture, the authors deployed a visual sequence-to-sequence model that combines global video features with fine-grained local visual cues from handshapes and mouth movements. To overcome the lack of gloss annotations, the method uses an automated sign-spotting technique on the raw video as a pre-training step.
The evaluation produced four key findings. First, pre-training the visual backbone using automated sign search and isolated sign datasets yielded an average 10% relative improvement over models pre-trained only on isolated signs, and sign-specific pre-training outperformed general action recognition pre-training by more than double across standard evaluation metrics. Second, incorporating specialized handshape and mouth features improved overall translation scores by approximately 5% relative to using global video features alone. Third, the full proposed approach outperformed prior baselines by roughly 15% across standard metrics, achieving a BLEU-4 score of 6.72 on the test set. Fourth, performance varied dramatically depending on sentence characteristics: the model achieved a 72.91 BLEU-4 score on frequently repeated short phrases, but only 4.09 on non-duplicate sentences, while also struggling significantly on clips containing fingerspelled words such as names and proper nouns.
These findings demonstrate that automated sign spotting and multi-region visual tracking provide a viable, cost-effective pathway to train translation systems without expensive gloss annotations. However, the low absolute performance on novel and complex sentences indicates that end-to-end sign language translation in open domains remains far below the maturity of spoken language translation, meaning automated systems cannot yet serve as reliable replacements for professional human interpreters in real-world settings.
Stakeholders and researchers should use the publicly available OpenASL dataset to benchmark future gloss-free translation systems. Development priorities must focus on handling unseen vocabulary, resolving fingerspelling for proper nouns, and transitioning from whole-word translation models to sub-word or character-aware architectures that can translate complex, spontaneous signing.
Confidence in these findings is solid regarding relative model improvements, supported by rigorous manual verification of test set alignments by professional ASL interpreters. Nonetheless, decision-makers must treat absolute performance cautiously: the dataset currently contains only one English reference sentence per video, lacks demographic ground-truth labels for signer background, and exhibits high error rates on spontaneous signing and fingerspelling.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
