Open-Domain Sign Language Translation Learned from Online Video

Bowen ShiDiane BrentariGregory ShakhnarovichKaren Livescu

article2022EMNLP120 citations

Presents OpenASL, the largest open-domain American Sign Language dataset from online video, alongside pre-training and multi-feature fusion techniques that substantially improve gloss-free translation in unconstrained environments.

Listen

Automatic sign language translation is critical for improving artificial intelligence accessibility for more than 430 million deaf and hard-of-hearing individuals worldwide. However, existing research has largely relied on small datasets recorded in controlled studio settings or narrow domains, such as weather forecasts. Furthermore, prevailing translation models heavily depend on intermediate gloss annotations—transliterations that directly transcribe sign language—which are costly, difficult to scale, and rarely available in real-world scenarios.

The article introduces OpenASL, a large-scale American Sign Language (ASL) dataset, and evaluates a gloss-free translation framework designed to handle continuous signing directly from real-world video into written English sentences.

To build the dataset, the authors collected 288 hours of online videos from YouTube news channels and community video blogs, covering over 200 signers across 98,417 video-sentence translation pairs. Because the source videos were self-generated with aligned English captions rather than interpreted broadcasts, sentence-level boundaries showed high temporal alignment without requiring full manual re-captioning. For the translation architecture, the authors deployed a visual sequence-to-sequence model that combines global video features with fine-grained local visual cues from handshapes and mouth movements. To overcome the lack of gloss annotations, the method uses an automated sign-spotting technique on the raw video as a pre-training step.

The evaluation produced four key findings. First, pre-training the visual backbone using automated sign search and isolated sign datasets yielded an average 10% relative improvement over models pre-trained only on isolated signs, and sign-specific pre-training outperformed general action recognition pre-training by more than double across standard evaluation metrics. Second, incorporating specialized handshape and mouth features improved overall translation scores by approximately 5% relative to using global video features alone. Third, the full proposed approach outperformed prior baselines by roughly 15% across standard metrics, achieving a BLEU-4 score of 6.72 on the test set. Fourth, performance varied dramatically depending on sentence characteristics: the model achieved a 72.91 BLEU-4 score on frequently repeated short phrases, but only 4.09 on non-duplicate sentences, while also struggling significantly on clips containing fingerspelled words such as names and proper nouns.

These findings demonstrate that automated sign spotting and multi-region visual tracking provide a viable, cost-effective pathway to train translation systems without expensive gloss annotations. However, the low absolute performance on novel and complex sentences indicates that end-to-end sign language translation in open domains remains far below the maturity of spoken language translation, meaning automated systems cannot yet serve as reliable replacements for professional human interpreters in real-world settings.

Stakeholders and researchers should use the publicly available OpenASL dataset to benchmark future gloss-free translation systems. Development priorities must focus on handling unseen vocabulary, resolving fingerspelling for proper nouns, and transitioning from whole-word translation models to sub-word or character-aware architectures that can translate complex, spontaneous signing.

Confidence in these findings is solid regarding relative model improvements, supported by rigorous manual verification of test set alignments by professional ASL interpreters. Nonetheless, decision-makers must treat absolute performance cautiously: the dataset currently contains only one English reference sentence per video, lacks demographic ground-truth labels for signer background, and exhibits high error rates on spontaneous signing and fingerspelling.

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for Open-Domain Sign Language Translation Learned from Online Video

Abstract

Existing work on sign language translation—that is, translation from sign language videos into sentences in a written language—has focused mainly on (1) data collected in a controlled environment or (2) data in a specific domain, which limits the applicability to real-world settings. In this paper, we introduce OpenASL, a large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube). OpenASL contains 288 hours of ASL videos in multiple domains from over 200 signers and is the largest publicly available ASL translation dataset to date. To tackle the challenges of sign language translation in realistic settings and without glosses, we propose a set of techniques including sign search as a pretext task for pre-training and fusion of mouthing and handshape features. The proposed techniques produce consistent and large improvements in translation quality, over baseline models based on prior work.

Table of Contents

  • 1 Introduction
  • 2.2 Methods for SLT
  • 2 Related Work
  • 2.1 Datasets for SLT
  • 2.3 Other related work
  • 3 The OpenASL Dataset
  • 4 Model and pre-training for gloss-free translation
  • 4.1 Sign spotting pre-training
  • 4.2 Hand and mouth ROI encoding
  • 4.3 Fusion and sequence modeling
  • 5 Experiments
  • 5.1 Setup
  • 5.2 Main Results
  • 5.3 Ablation Study
  • 5.4 Analysis
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Existing datasets
  • A.3 Sign Search
  • A.4 Implementation details
  • A.2 Instructions for meta annotation
  • A.5 Baseline performance on Phoenix-14T
  • A.6 Does fine-tuning the visual encoder help?
  • A.7 Which pre-training data to use?
  • A.8 Fingerspelling vs. non-fingerspelling
  • A.9 Translation examples
  • A.10 Visualization

Knowls

  1. Knowl 1 — OpenASL provides a large, diverse ASL–English translation corpus

    data/table

    OpenASL contains 288 hours of American Sign Language (ASL) video paired with English translations, comprising 98,417 video–sentence pairs and 33,549 distinct words. The videos come mainly from online sources: news signing from TheDailyMoth and Sign1News, and National Association of the Deaf (NAD) videos that include VLOGs, announcements, tips, and conversations. The dataset includes more than 200 signers (approximately 220; some identities are unknown) and spans multiple domains and real-world visual conditions. The authors split subtitle transcripts into sentences to form the pairs, then randomly selected 966 pairs for validation and 975 for test; the remaining pairs constitute training data.

  2. Knowl 2 — Professional review found the subtitles and sentence timing generally reliable

    empirical result

    For OpenASL’s validation and test sets, professional ASL interpreters reviewed each clip’s English translation and time boundaries, correcting them when needed and also supplying gloss sequences. They could consult the full video for context. In a manually checked subset, subtitle sentence boundaries usually differed from signing boundaries by less than 2 seconds, while fewer than 5% of the shifts exceeded 2 seconds. Comparing original subtitle translations with corrected translations yielded 81.0 BLEU-4. On the basis of this agreement and timing quality, the authors did not proofread the training-set translations.

  3. Knowl 3 — Weakly supervised sign search supplies in-domain pretraining examples

    algorithm

    The sign-search procedure uses each training pair of ASL video and English sentence to find candidate lexical signs and fingerspelled words without gloss annotations. It returns video intervals paired with words from the sentence.

    1. For lexical signs, slide a 32-frame window over the video at a stride of 8 frames. Apply an isolated-sign recognizer to each window. Among words that occur in both the English sentence and the recognizer’s vocabulary, select the word with the highest predicted probability; retain that interval–word pair if the probability exceeds 0.6.
    2. For fingerspelling, run a fingerspelling detector over the video and retain proposals with detector confidence above 0.5. Apply a fingerspelling recognizer to each proposal, compare its spelling hypothesis with each word in the English sentence using letter accuracy, and retain the best-matching interval–word pair if letter accuracy exceeds 0.2.
    3. Combine the retained lexical and fingerspelled pairs. The procedure found 32,602 signs in the translation data. The authors use the spotted examples together with WLASL isolated-sign data to train the I3D visual backbone: 50 epochs of SGD, batch size 8, learning rate 0.01, momentum 0.9, and halving the learning rate at epochs 20 and 40.

    The lexical recognizer is an I3D model trained for isolated ASL sign recognition; the fingerspelling detector proposes intervals and a separate fingerspelling recognizer produces spelling hypotheses. The sentence text provides the weak supervision used to associate candidate intervals with words.

  4. Knowl 4 — The translation model fuses global, handshape, and mouthing representations

    model/method

    The gloss-free system maps a video clip to English words using a global visual stream and two local streams. An I3D backbone, initialized with Kinetics and trained on ASL sign data including the spotted signs, extracts global video features. A fingerspelling recognizer supplies hand-region features for both hands, which are concatenated; an external English lip-reading model, AV-HuBERT, extracts features from the signer’s mouth region. The visual backbones are frozen while training the translation model.

    Each of the three feature sequences—global, hand, and mouth—is encoded by its own two-layer Transformer encoder. A shared two-layer Transformer decoder uses separate cross-attention over the three encoded sequences at each decoding step; it concatenates the resulting context vectors and passes them through a feedforward layer to predict the next word. The encoder and decoder use 512-dimensional hidden states and 2,048-dimensional feedforward layers. The output vocabulary contains 21,475 words. Translation training uses Adam for 14,000 iterations with batch size 64; the learning rate rises linearly to 0.001 over the first 2,000 iterations and then decays to zero. Decoding uses beam search, with beam width and length penalty selected on validation data.

  5. Knowl 5 — The full system outperforms gloss-free baseline models on OpenASL

    empirical result

    On OpenASL, the proposed model combines sign-search pretraining with handshape and mouthing features. It is compared with a Conv-GRU model using an ImageNet-pretrained AlexNet backbone and an I3D-Transformer using global features with WLASL sign pretraining. The proposed system scores highest on every reported development and test metric. Relative to the I3D-Transformer, the paper reports average gains of about 15% in ROUGE and BLEU scores.

    Model Split ROUGE BLEU-1 BLEU-2 BLEU-3 BLEU-4 BLEURT
    Conv-GRU Dev 16.25 16.72 8.95 6.31 4.82 25.36
    Conv-GRU Test 16.10 16.11 8.85 6.18 4.58 25.65
    I3D-Transformer Dev 18.88 18.26 10.26 7.17 5.60 29.17
    I3D-Transformer Test 18.64 18.31 10.15 7.19 5.66 28.82
    Proposed model Dev 20.43 20.10 11.81 8.43 6.57 31.22
    Proposed model Test 21.02 20.92 12.08 8.59 6.72 31.09

    These scores establish improvement over the two baselines, but the absolute translation scores remain low.

  6. Knowl 6 — Adding spotted signs improves translation over isolated-sign pretraining alone

    empirical result

    On the development set, the authors compared global-feature I3D-Transformer models without local hand or mouth features. The model pretrained on WLASL alone scored ROUGE 18.88, BLEU-1 18.26, BLEU-2 10.26, BLEU-3 7.17, and BLEU-4 5.60. Adding spotted lexical and fingerspelled signs from the translation training data changed the scores to ROUGE 19.65, BLEU-1 19.72, BLEU-2 11.18, BLEU-3 8.56, and BLEU-4 6.51. The paper describes this as an average relative improvement of approximately 10% across metrics and attributes the gains to adapting the visual backbone to continuous signing and real-world video conditions absent from isolated-sign data.

  7. Knowl 7 — Handshape and mouthing features add gains beyond global video features

    empirical result

    On the development set, the authors compared models with and without local features; both used spotted-sign pretraining. The global-only model scored ROUGE 19.65, BLEU-1 19.72, BLEU-2 11.08, BLEU-3 8.06, and BLEU-4 6.30. Adding both handshape and mouthing features raised these scores to ROUGE 20.43, BLEU-1 20.10, BLEU-2 11.81, BLEU-3 8.43, and BLEU-4 6.57. The reported improvement is around 5% across metrics, with relatively larger gains on lower-order BLEU.

  8. Knowl 8 — Fine-tuning the visual backbone lowers most translation metrics

    empirical result

    In a development-set comparison of I3D visual-backbone training strategies, freezing the backbone produced ROUGE 18.88, BLEU-1 18.26, BLEU-2 10.26, BLEU-3 7.17, and BLEU-4 5.60. Fine-tuning produced ROUGE 18.91, BLEU-1 16.95, BLEU-2 9.12, BLEU-3 5.87, and BLEU-4 4.38. Thus fine-tuning slightly raised ROUGE but reduced all four BLEU scores, especially the higher-order scores. The authors suggest that the available paired data may be insufficient for end-to-end visual-backbone adaptation and that English sentence supervision may be too weak to train the visual encoder effectively.

  9. Knowl 9 — Repeated sentences strongly affect reported evaluation performance

    empirical result

    Some English sentences occur in both the OpenASL training and evaluation sets. The paper reports duplicates in 10.9% of development examples (105 of 967) and 10.6% of test examples (103 of 976), often common phrases such as greetings and thanks. The proposed model’s BLEU-4 was 72.91 on the duplicate subset but only 4.09 on the non-duplicate subset. The authors note that duplicate examples are often short and may be easy to memorize, so overall scores should be interpreted in light of this train–evaluation sentence overlap.

  10. Knowl 10 — OpenASL and the model have important coverage and performance limitations

    limitation

    OpenASL is large relative to publicly available ASL translation datasets but remains small compared with written-language translation corpora, and it provides only one English translation per video. The dataset has no ground-truth labels for signer gender, race, or handedness, so the authors cannot establish the distribution of those characteristics. The sign-search method depends on isolated-sign data, and the whole-word translation vocabulary cannot adequately handle words unseen during training or ASL signs without an English word equivalent. In a development-set analysis, 54.7% of clips contained at least one fingerspelled word; BLEU-4 was 6.33 on clips with fingerspelling and 7.74 on clips without it. The paper identifies rare or long fingerspelled words, particularly proper nouns, as a recurring difficulty. Overall translation quality remains low, and the system is not suitable as a replacement for human interpreters in real-world settings.

Coverage note — The qualitative translation and sign-spotting examples, and the news-versus-VLOG subgroup comparison, were omitted because they are illustrative or offer only qualitative explanations rather than additional independently reported quantitative findings.

References

  1. 1.Nikolaos Adaloglou, Theocharis Chatzis, Ilias Papastratis, Andreas Stergioulas, Georgios Papadopoulos, Vassia Zacharopoulou, George Xydopoulos, Klimis Antzakas, Dimitris Papazachariou, and Petros Daras. 2021. A comprehensive study on deep learning-based methods for sign language recognition. IEEE Transactions on Multimedia, PP:1–1.
  2. 2.Samuel Albanie, Gül Varol, Liliane Momeni, Triantafyllos Afouras, Joon Son Chung, Neil Fox, and Andrew Zisserman. 2020. BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues. In ECCV.
  3. 3.Samuel Albanie, Gül Varol, Liliane Momeni, Hannah Bull, Triantafyllos Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, and Andrew Zisserman. 2021. BOBSL: BBC-Oxford British Sign Language Dataset.
  4. 4.Danielle Bragg, Oscar Koller, Mary Bellard, Larwan Berke, Patrick Boudreault, Annelies Braffort, Naomi K. Caselli, Matt Huenerfauth, Hernisa Kacorri, Tessa Verhoef, Christian Vogler, and Meredith Ringel Morris. 2019. Sign language recognition, generation, and translation: An interdisciplinary perspective. The 21st International ACM SIGACCESS Conference on Computers and Accessibility.
  5. 5.Patrick Buehler, Andrew Zisserman, and Mark Everingham. 2009. Learning sign language by watching tv (using weakly aligned subtitles). In CVPR, pages 2961–2968.
  6. 6.Hannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie, Liliane Momeni, and Andrew Zisserman. 2021. Aligning subtitles in sign language videos. In CVPR.
  7. 7.N. C. Camgoz, B. Saunders, G. Rochette, M. Giovanelli, G. Inches, R. Nachtrab-Ribback, and R. Bowden. 2021. Content4all open research sign language translation datasets. ArXiv, abs/2105.02351.
  8. 8.Necati Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden. 2020a. Multi-channel transformers for multi-articulatory sign language translation. In ECCV, pages 301–319.
  9. 9.Necati Cihan Camgoz, Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden. 2018. Neural sign language translation. In CVPR, pages 7784–7793.
  10. 10.Necati Cihan Camgoz, Oscar Koller, Simon Hadfield, and Richard Bowden. 2020b. Sign language transformers: Joint end-to-end sign language recognition and translation. In CVPR.
  11. 11.João Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. CVPR, pages 4724–4733.
  12. 12.Yutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu, and Stephen Lin. 2022. A simple multi-modality transfer learning baseline for sign language translation. In CVPR.
  13. 13.Philippe Dreuw, David Rybach, Thomas Deselaers, Morteza Zahedi, and Hermann Ney. 2007. Speech recognition techniques for a sign language recognition system. In Interspeech, volume 1, pages 2513–2516.
  14. 14.Amanda Duarte, Shruti Palaskar, Lucas Ventura, Deepti Ghadiyaram, Kenneth DeHaan, Florian Metze, Jordi Torres, and Xavier Giro-i Nieto. 2021. How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign Language. In CVPR.
  15. 15.Gunnar Farnebäck. 2003. Two-frame motion estimation based on polynomial expansion. In SCIA.
  16. 16.Shiwei Gan, Yafeng Yin, Zhiwei Jiang, Linfu Xie, and Sanglu Lu. 2021. Skeleton-aware neural sign language translation. Proceedings of the 29th ACM International Conference on Multimedia.
  17. 17.Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In ICML.
  18. 18.Davis E. King. 2009. Dlib-ml: A machine learning toolkit. J. Mach. Learn. Res., 10:1755–1758.
  19. 19.Diederik P. Kingma and Jimmy Ba. 2015. ADAM: A method for stochastic optimization. CoRR, abs/1412.6980.
  20. 20.S. Ko, C. Kim, H. Jung, and C. Cho. 2019. Neural sign language translation based on human keypoint estimation. Applied Sciences, 9.
  21. 21.Oscar Koller, Necati Cihan Camgoz, Hermann Ney, and Richard Bowden. 2020. Weakly supervised learning with multi-stream cnn-lstm-hmms to discover sequential parallelism in sign language videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(9):2306–2320.
  22. 22.Dongxu Li, Cristian Rodriguez-Opazo, Xin Yu, and Hongdong Li. 2020a. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In WACV, pages 1448–1458.
  23. 23.Dongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang, Ben Swift, Hanna Suominen, and Hongdong Li. 2020b. TSPNet: Hierarchical feature learning via temporal semantic pyramid for sign language translation. In NeurIPS, volume abs/2010.05468.
  24. 24.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In ACL.
  25. 25.Liliane Momeni, Gül Varol, Samuel Albanie, Triantafyllos Afouras, and Andrew Zisserman. 2020. Watch, read and lookup: learning to spot signs from multiple supervisors. In ACCV.
  26. 26.Marie A. Nadolske and Rachel Rosenstock. 2008. Occurrence of mouthings in American Sign Language: A preliminary study, pages 35–62. De Gruyter Mouton.
  27. 27.Alptekin Orbay and Lale Akarun. 2020. Neural sign language translation by learning tokenization. In IEEE International Conference on Automatic Face and Gesture Recognition, pages 222–228.
  28. 28.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In ACL.
  29. 29.Razieh Rastgoo, Kourosh Kiani, and Sergio Escalera. 2021. Sign language recognition: A deep survey. Expert Systems with Applications, 164:113794.
  30. 30.Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020. BLEURT: Learning robust metrics for text generation. In ACL.
  31. 31.Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. 2021. Fingerspelling detection in american sign language. In CVPR, pages 4164–4173.
  32. 32.Bowen Shi, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. 2022a. Searching for fingerspelled content in american sign language. In ACL.
  33. 33.Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. 2022b. Learning audio-visual speech representation by masked multimodal cluster prediction. In ICLR.
  34. 34.Bowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari, Greg Shakhnarovich, and Karen Livescu. 2019. Fingerspelling recognition in the wild with iterative visual attention. In ICCV, pages 5399–5408.
  35. 35.Bowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Jonathan Michaux, Diane Brentari, Gregory Shakhnarovich, and Karen Livescu. 2018. American sign language fingerspelling recognition in the wild. In SLT, pages 145–152.
  36. 36.Dimitar Shterionov. 2021. Proceedings of the 1st international workshop on automatic translation for signed and spoken languages (at4ssl). In Proceedings of the 1st International Workshop on Automatic Translation for Signed and Spoken Languages (AT4SSL).
  37. 37.Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. 2019. Deep high-resolution representation learning for human pose estimation. In CVPR, pages 5686–5696.
  38. 38.Gül Varol, Liliane Momeni, Samuel Albanie, Triantafyllos Afouras, and Andrew Zisserman. 2021. Read and attend: Temporal localisation in sign language videos. In CVPR, pages 16852–16861.
  39. 39.Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS.
  40. 40.R. B. Wilbur et al. 2006. Purdue RVL-SLLL American Sign Language Database. Technical report, School of Electrical and Computer Engineering, Purdue University.
  41. 41.K. Yin, A. Moryossef, J. Hochgesang, Y. Goldberg, and M. Alikhani. 2021. Including signed languages in natural language processing. In ACL.
  42. 42.Kayo Yin and Jesse Read. 2020. Better sign language translation with STMC-Transformer. In COLING.
  43. 43.Hao Zhou, Wen gang Zhou, Weizhen Qi, Junfu Pu, and Houqiang Li. 2021. Improving sign language translation with monolingual data by sign back-translation. In CVPR, pages 1316–1325.

Citation

MLA
Shi, B., et al. “Open-Domain Sign Language Translation Learned from Online Video”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 6365–79, https://doi.org/10.18653/v1/2022.emnlp-main.427.
APA
Shi, B., Brentari, D., Shakhnarovich, G., & Livescu, K. (2022). Open-Domain Sign Language Translation Learned from Online Video. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6365–6379. https://doi.org/10.18653/v1/2022.emnlp-main.427
Chicago
Shi, B., D. Brentari, G. Shakhnarovich, and K. Livescu. 2022. “Open-Domain Sign Language Translation Learned from Online Video”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 6365–79. https://doi.org/10.18653/v1/2022.emnlp-main.427.
Harvard
Shi, B. et al. (2022) “Open-Domain Sign Language Translation Learned from Online Video”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 6365–6379. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.427.
Vancouver
1. Shi B, Brentari D, Shakhnarovich G, Livescu K (2022) Open-Domain Sign Language Translation Learned from Online Video. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 6365–6379

BibTeX

@inproceedings{shi-etal-2022-open,
    title = "Open-Domain Sign Language Translation Learned from Online Video",
    author = "Shi, Bowen  and
      Brentari, Diane  and
      Shakhnarovich, Gregory  and
      Livescu, Karen",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.427/",
    doi = "10.18653/v1/2022.emnlp-main.427",
    pages = "6365--6379"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/