Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

Peng QiYuhao ZhangYuhui ZhangJason BoltonChristopher D. Manning

article2020ACL2,183 citations

Presents Stanza, an open-source Python package that provides a language-agnostic, fully neural text analysis pipeline across 66 languages for core linguistic tasks from tokenization to named entity recognition.

Listen

Modern natural language processing systems frequently face operational bottlenecks when handling global, multilingual text. Most established software packages support only a handful of dominant languages, rely on older statistical models that limit accuracy, or require pre-processed inputs rather than handling raw text end-to-end. As organizations increasingly need to extract structured insights from diverse, international communication channels, the lack of accurate, unified, and broad-coverage language tools poses a significant technical barrier.

The article evaluates whether an open-source, fully neural processing framework named Stanza can deliver state-of-the-art linguistic analysis across a wide range of human languages directly from raw text. It also assesses the integration of a dedicated Python interface to connect modern data workflows with the extensive capabilities of Stanford's existing Java-based processing suite.

To demonstrate this, the authors constructed a modular deep-learning pipeline covering core text analysis tasks, from basic sentence segmentation to complex syntactic parsing and entity recognition. They trained and benchmarked the system across 112 standardized evaluation datasets, encompassing 66 languages from diverse language families, and compared its performance and speed directly against leading alternatives.

The analysis yielded several key findings. First, Stanza achieved state-of-the-art or highly competitive accuracy across all 66 evaluated languages, consistently outperforming established alternatives on standard syntactic analysis benchmarks. Second, in identifying names and entities across eight major languages, Stanza matched or exceeded top-tier specialized tools while reducing model storage size by up to 75 percent. Third, the system demonstrated strong versatility by converting raw text into detailed grammatical structures without requiring external pre-processing tools. Finally, runtime evaluations confirmed that while the neural architecture requires significantly more computation time than purely speed-focused tools on standard processors, it achieves competitive processing speeds when accelerated by graphical processing units.

These results show that organizations no longer need to sacrifice linguistic depth or broad language coverage when building text analysis pipelines. Deploying a unified, highly accurate architecture reduces the risk of errors cascading into downstream business intelligence or automated decision systems. However, teams must weigh the trade-off between higher computational overhead and superior analytical accuracy based on their specific operational latency requirements.

Organizations handling multilingual text should consider adopting or piloting this framework, particularly where GPU resources are available to offset processing demands. For future developments, the authors highlight the need to expand pre-trained models across blended text genres, establish an open community repository for model sharing, and explore compression techniques to further improve processing speed without degrading accuracy.

Decision-makers should note that current models are largely trained on single-domain benchmarks, meaning performance may vary when applied to niche or out-of-domain text. Nonetheless, the extensive empirical testing across over one hundred datasets provides high confidence in the framework's baseline accuracy and multilingual robustness.

  • Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). Read this paper next to explore how massively multilingual text-to-text transformer models extend language coverage beyond Stanza's 66 languages.
  • Paper: Qwen2.5 Technical Report, Qwen et al. (2024). This report continues the trajectory of multilingual language technology by showcasing how modern large language models handle extreme context windows.
Cover for Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

Abstract

We introduce Stanza, an open-source Python natural language processing toolkit supporting 66 human languages. Compared to existing widely used toolkits, Stanza features a language-agnostic fully neural pipeline for text analysis, including tokenization, multi-word token expansion, lemmatization, part-of-speech and morphological feature tagging, dependency parsing, and named entity recognition. We have trained Stanza on a total of 112 datasets, including the Universal Dependencies treebanks and other multilingual corpora, and show that the same neural architecture generalizes well and achieves competitive performance on all languages tested. Additionally, Stanza includes a native Python interface to the widely used Java Stanford CoreNLP software, which further extends its functionality to cover other tasks such as coreference resolution and relation extraction. Source code, documentation, and pretrained models for 66 languages are available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 System Design and Architecture
  • 2.1 Neural Multilingual NLP Pipeline
  • 2.2 CoreNLP Client
  • 3 System Usage
  • 3.1 Neural Pipeline Interface
  • 3.2 CoreNLP Client Interface
  • 3.3 Interactive Web-based Demo
  • 3.4 Training Pipeline Models
  • 4 Performance Evaluation
  • 5 Conclusion and Future Work
  • References

Knowls

  1. Knowl 1 — Fully Neural NLP Pipeline Architecture in Stanza

    model/method

    Stanza provides a fully neural, language-agnostic natural language processing pipeline that converts raw text into linguistic annotations across six sequential modular processors:

    1. Tokenization and Sentence Splitting (tokenize): Jointly segments raw text into sentences and tokens, while identifying multi-word tokens (MWTs).
    2. Multi-Word Token Expansion (mwt): Expands composite tokens into their constituent syntactic words.
    3. Part-of-Speech and Morphological Feature Tagging (pos): Assigns Universal POS (UPOS), treebank-specific POS (XPOS), and Universal Morphological Features (UFeats).
    4. Lemmatization (lemma): Infers the canonical dictionary form for each word.
    5. Dependency Parsing (depparse): Produces syntactic dependency trees linking words to their syntactic heads with dependency relations.
    6. Named Entity Recognition (ner): Identifies entity spans and classifies them into categories (e.g., person, organization).

    Annotations are exposed as native Python objects organized hierarchically: a Document contains a sequence of Sentence objects, which contain Token objects (representing contiguous spans of characters in the input text) and Word objects (representing syntactic words). In languages without multi-word tokens, tokens and words have a one-to-one correspondence; where multi-word tokens occur, one token maps to multiple syntactic words.

  2. Knowl 2 — Character-Level Joint Tokenization, Sentence Splitting, and Multi-Word Token Identification

    model/method

    Stanza unifies tokenization, sentence boundary segmentation, and multi-word token (MWT) detection from raw character sequences into a single sequence tagging module.

    Given raw text as an input character sequence C=(c1,c2,,cN)C = (c_1, c_2, \dots, c_N), the neural model performs character-level classification to predict whether each character cic_i represents:

    • The end of a regular token
    • The end of a sentence
    • The end of an MWT span

    Jointly predicting multi-word token boundaries at the character level enables context-sensitive segmentation in languages where identical character sequences can function either as single words or as contractions that require expansion (for example, in French, where des can act as a single word or as the contraction of de and les).

  3. Knowl 3 — Hybrid Lexicon and Sequence-to-Sequence Modeling for MWT Expansion and Lemmatization

    model/method

    Stanza handles morphological expansions (multi-word token expansion and lemmatization) via hybrid ensembles combining frequency dictionaries with neural sequence-to-sequence (seq2seq) models:

    • Multi-word Token Expansion: Identified multi-word tokens are expanded into constituent syntactic words using an ensemble of a frequency lexicon extracted from training data and a neural character-level seq2seq model. Frequently occurring contractions are reliably resolved via dictionary lookup, while rare or unseen forms are generated statistically by the seq2seq model.
    • Lemmatization: Each word is mapped to its base lemma (e.g., did \to do) using an ensemble of a dictionary-based lemmatizer and a neural seq2seq lemmatizer. The seq2seq encoder features an auxiliary classifier that predicts shortcut transformations (such as identity copy or lowercasing), ensuring robust behavior on long inputs and out-of-vocabulary strings like URLs.
  4. Knowl 4 — Part-of-Speech and Morphological Tagging via Bi-LSTM with Biaffine Conditioning

    model/method

    Stanza predicts Universal Part-of-Speech tags (UPOS), treebank-specific Part-of-Speech tags (XPOS), and Universal Morphological Features (UFeats, e.g., gender, number, tense, person) for each word in a sentence using a Bidirectional Long Short-Term Memory (Bi-LSTM) network.

    To ensure consistency across the three prediction spaces, Stanza conditions the predictions of XPOS and UFeats on the UPOS representations using a deep biaffine scoring mechanism. The biaffine classifier computes transformation matrices that project the hidden representation of the word alongside representations derived from UPOS scoring to jointly assign compatible XPOS and UFeats tags.

  5. Knowl 5 — Neural Dependency Parsing with Linearization and Distance Prediction

    model/method

    Stanza performs syntactic dependency parsing using a graph-based deep biaffine neural dependency parser over Bi-LSTM word representations. Each word is assigned a directed syntactic head (either another word in the sentence or an artificial root symbol) and a dependency relation label.

    To improve parsing accuracy across diverse linguistic typologies, the biaffine architecture incorporates two auxiliary linguistically motivated feature objectives:

    1. Linearization Order Prediction: An auxiliary classifier predicting the expected linear word order between head and dependent tokens in the target language.
    2. Distance Prediction: An auxiliary classifier predicting the typical linear token distance between the dependent and its syntactic head.
  6. Knowl 6 — Contextualized Character-Word Representation Architecture for Named Entity Recognition

    model/method

    Stanza implements Named Entity Recognition (NER) using a contextualized sequence labeling architecture:

    1. Pretrained Character Language Models: Forward and backward character-level Long Short-Term Memory (LSTM) language models are pretrained on large multilingual text corpora (such as Common Crawl, Wikipedia, WMT news data, Google One Billion Word, and Chinese Gigaword).
    2. Representation Concatenation: At tagging time, contextualized character representations extracted from the final character position of each word (from both the forward and backward character LSTMs) are concatenated with static word embeddings (such as word2vec or fastText).
    3. Sequence Tagger: The combined representation is fed into a 1-layer Bidirectional LSTM.
    4. Inference: Entity span tags are decoded globally using a linear-chain Conditional Random Field (CRF) decoder.
  7. Knowl 7 — Transparent Local Server Client for Stanford CoreNLP

    model/method

    Stanza provides a native Python client interface (CoreNLPClient) to the Java Stanford CoreNLP toolkit with the following operational mechanisms:

    • Automated Lifecycle Management: Instantiating CoreNLPClient automatically launches the Java CoreNLP server as a local background child process.
    • RESTful API and Serialization: Client-server communication occurs over HTTP RESTful endpoints. Linguistic annotations are serialized using Protocol Buffers (or alternatively JSON/XML) and mapped into native Python objects.
    • Fault Tolerance: The client conducts periodic health checks against the local server and automatically restarts the background Java process if a timeout or crash is detected.
    • Extended Linguistic Tasks: Enables Python access to Stanford CoreNLP annotators outside Stanza's native neural pipeline, such as coreference resolution and relation extraction.
  8. Knowl 8 — Stanza Parsing and Tagging Performance on Universal Dependencies v2.5

    data/table

    Stanza (v1.0) was evaluated on the Universal Dependencies (UD) v2.5 test treebanks, covering 100 treebanks across 66 languages. Performance was macro-averaged over all 100 treebanks and compared against UDPipe (v1.2) and spaCy (v2.2) on five major language treebanks using the official CoNLL 2018 UD Shared Task evaluation script (F1F_1 scores):

    Treebank System Tokens Sents. Words UPOS XPOS UFeats Lemmas UAS LAS
    Overall (100 treebanks) Stanza 99.09 86.05 98.63 92.49 91.80 89.93 92.78 80.45 75.68
    Arabic-PADT Stanza 99.98 80.43 97.88 94.89 91.75 91.86 93.27 83.27 79.33
    UDPipe 99.98 82.09 94.58 90.36 84.00 84.16 88.46 72.67 68.14
    Chinese-GSD Stanza 92.83 98.80 92.83 89.12 88.93 92.11 92.83 72.88 69.82
    UDPipe 90.27 99.10 90.27 84.13 84.04 89.05 90.26 61.60 57.81
    English-EWT Stanza 99.01 81.13 99.01 95.40 95.12 96.11 97.21 86.22 83.59
    UDPipe 98.90 77.40 98.90 93.26 92.75 94.23 95.45 80.22 77.03
    spaCy 97.30 61.19 97.30 86.72 90.83 87.05
    French-GSD Stanza 99.68 94.92 99.48 97.30 96.72 97.64 91.38 89.05
    UDPipe 99.68 93.59 98.81 95.85 95.55 96.61 87.14 84.26
    spaCy 98.34 77.30 94.15 86.82 87.29 67.46 60.60
    Spanish-AnCora Stanza 99.98 99.07 99.98 98.78 98.67 98.59 99.19 92.21 90.01
    UDPipe 99.97 98.32 99.95 98.32 98.13 98.13 98.48 88.22 85.10
    spaCy 99.47 97.59 98.95 94.04 79.63 86.63 84.13

    Stanza achieves higher labeled attachment scores (LAS) and unlabeled attachment scores (UAS) than UDPipe and spaCy across all evaluated treebanks (e.g., 83.5983.59 LAS on English-EWT vs 77.0377.03 for UDPipe; 89.0589.05 LAS on French-GSD vs 84.2684.26 for UDPipe and 60.6060.60 for spaCy).

  9. Knowl 9 — Multilingual Named Entity Recognition Performance Comparison

    data/table

    Stanza (v1.0) was evaluated on 12 named entity recognition datasets across 8 human languages, comparing entity micro-averaged test F1F_1 scores against FLAIR (v0.4.5) and spaCy (v2.2):

    Language Corpus # Types Stanza FLAIR spaCy
    Arabic AQMAR 4 74.3 74.0
    Chinese OntoNotes 18 79.2
    Dutch CoNLL02 4 89.2 90.3 73.8
    WikiNER 4 94.8 94.8 90.9
    English CoNLL03 4 92.1 92.7 81.0
    OntoNotes 18 88.8 89.0 85.4
    French WikiNER 4 92.9 92.5 88.8
    German CoNLL03 4 81.9 82.5 63.9
    GermEval14 4 85.2 85.4 68.4
    Russian WikiNER 4 92.9
    Spanish CoNLL02 4 88.1 87.3 77.5
    AnCora 4 88.6 88.4 76.1

    Stanza performs comparably to or better than FLAIR across all 12 benchmarks while compressing pretrained model disk footprints by up to 75% relative to FLAIR. Stanza outperforms spaCy across all corpora, achieving improvements ranging from +3.4+3.4 to +18.0+18.0 F1F_1 points.

  10. Knowl 10 — Annotation Runtime Relative to spaCy CPU

    data/table

    Inference runtimes for Stanza, UDPipe, and FLAIR were benchmarked relative to CPU-based spaCy on the English Universal Dependencies EWT treebank (UD task) and the OntoNotes NER test set (NER task). GPU measurements used a single NVIDIA Titan RTX card. As a reference baseline, spaCy on CPU processes 8,140 tokens per second for UD and 5,912 tokens per second for NER:

    Stanza UDPipe FLAIR
    Task CPU GPU CPU CPU GPU
    UD 10.3×10.3\times 3.22×3.22\times 4.30×4.30\times
    NER 17.7×17.7\times 1.08×1.08\times 51.8×51.8\times 1.17×1.17\times

    Stanza's deep neural architecture requires more processing time than spaCy on CPU (10.3×10.3\times for UD and 17.7×17.7\times for NER). However, with GPU acceleration, Stanza processes NER at 1.08×1.08\times the runtime of spaCy CPU, while running nearly 3×3\times faster than FLAIR on CPU (17.7×17.7\times vs 51.8×51.8\times relative to spaCy).

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Alan Akbik, Tanja Bergmann, Duncan Blythe, Kashif Rasul, Stefan Schweter, and Roland Vollgraf. 2019. FLAIR: An easy-to-use framework for state-of-the-art NLP. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations). Association for Computational Linguistics.
  2. 2.Alan Akbik, Duncan Blythe, and Roland Vollgraf. 2018. Contextual string embeddings for sequence labeling. In Proceedings of the 27th International Conference on Computational Linguistics. Association for Computational Linguistics.
  3. 3.Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019. Findings of the 2019 conference on machine translation (WMT19). In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1). Association for Computational Linguistics.
  4. 4.Darina Benikova, Chris Biemann, and Marc Reznicek. 2014. NoSta-D named entity annotation for German: Guidelines and dataset. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14).
  5. 5.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5.
  6. 6.Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. Technical report, Google.
  7. 7.Timothy Dozat and Christopher D. Manning. 2017. Deep biaffine attention for neural dependency parsing. In International Conference on Learning Representations (ICLR).
  8. 8.Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations.
  9. 9.Behrang Mohit, Nathan Schneider, Rishav Bhowmick, Kemal Oflazer, and Noah A Smith. 2012. Recall-oriented learning of named entities in Arabic Wikipedia. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics.
  10. 10.Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020. Universal dependencies v2: An evergrowing multilingual treebank collection. In Proceedings of the Twelfth International Conference on Language Resources and Evaluation (LREC’20).
  11. 11.Joel Nothman, Nicky Ringland, Will Radford, Tara Murphy, and James R Curran. 2013. Learning multilingual named entity recognition from Wikipedia. Artificial Intelligence, 194:151–175.
  12. 12.Peng Qi, Timothy Dozat, Yuhao Zhang, and Christopher D. Manning. 2018. Universal dependency parsing from scratch. In Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Association for Computational Linguistics.
  13. 13.Milan Straka. 2018. UDPipe 2.0 prototype at CoNLL 2018 UD shared task. In Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Association for Computational Linguistics.
  14. 14.Mariona Taulé, M. Antònia Martí, and Marta Recasens. 2008. AnCora: Multilevel annotated corpora for Catalan and Spanish. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08). European Language Resources Association (ELRA).
  15. 15.Erik F. Tjong Kim Sang. 2002. Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition. In COLING-02: The 6th Conference on Natural Language Learning 2002 (CoNLL-2002).
  16. 16.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003.
  17. 17.Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, et al. 2013. OntoNotes release 5.0. Linguistic Data Consortium.
  18. 18.Daniel Zeman, Jan Hajic, Martin Popel, Martin Potthast, Milan Straka, Filip Ginter, Joakim Nivre, and Slav Petrov. 2018. CoNLL 2018 shared task: Multilingual parsing from raw text to universal dependencies. In Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Association for Computational Linguistics.
  19. 19.Daniel Zeman, Joakim Nivre, Mitchell Abrams, Noëmi Aepli, Željko Agic, Lars Ahrenberg, Gabrielė Aleksandravičiūtė, Lene Antonsen, Katya Aplonova, Maria Jesus Aranzabe, Gashaw Arutie, Masayuki Asahara, Luma Ateyah, Mohammed Attia, Aitziber Atutxa, Liesbeth Augustinus, Elena Badmaeva, Miguel Ballesteros, Esha Banerjee, Sebastian Bank, Verginica Barbu Mititelu, Victoria Basmov, Colin Batchelor, John Bauer, Sandra Bellato, Kepa Bengoetxea, Yevgeni Berzak, Irshad Ahmad Bhat, Riyaz Ahmad Bhat, Erica Biagetti, Eckhard Bick, Agne Bielinskienė, Rogier Blokland, Victoria Bobicev, Loïc Boizou, Emanuel Borges Völker, Carl Börstell, Cristina Bosco, Gosse Bouma, Sam Bowman, Adriane Boyd, Kristina Brokaitė, Aljoscha Burchardt, Marie Candito, Bernard Caron, Gauthier Caron, Tatiana Cavalcanti, Gülşen Cebiroğlu Eryiğit, Flavio Massimiliano Cecchini, Giuseppe G. A. Celano, Slavomír Céplö, Savaş Çetin, Fabricio Chalub, Jinho Choi, Yongseok Cho, Jayeol Chun, Alessandra T. Cignarella, Silvie Cinková, Aurélie Collomb, Çağrı Çöltekin, Miriam Connor, Marine Courtin, Elizabeth Davidson, Marie-Catherine de Marneffe, Valeria de Paiva, Elvis de Souza, Arantza Diaz de Ilarraza, Carly Dickerson, Bamba Dione, Peter Dirix, Kaja Dobrovoljc, Timothy Dozat, Kira Droganova, Puneet Dwivedi, Hanne Eckhoff, Marhaba Eli, Ali Elkahky, Binyam Ephrem, Olga Erina, Tomaž Erjavec, Aline Etienne, Wograine Evelyn, Richárd Farkas, Hector Fernandez Alcalde, Jennifer Foster, Cláudia Freitas, Kazunori Fujita, Katarína Gajdošová, Daniel Galbraith, Marcos Garcia, Moa Gärdenfors, Sebastian Garza, Kim Gerdes, Filip Ginter, Iakes Goenaga, Koldo Gojenola, Memduh Gökırmak, Yoav Goldberg, Xavier Gómez Guinovart, Berta González Saavedra, Bernadeta Griciūtė, Matias Grioni, Normunds Gruz̄ı̄tis, Bruno Guillaume, Céline Guillot-Barbance, Nizar Habash, Jan Hajic, Jan Hajic jr., Mika Hämäläinen, Linh Hà Mỹ, Na-Rae Han, Kim Harris, Dag Haug, Johannes Heinecke, Felix Hennig, Barbora Hladká, Jaroslava Hlavácová, Florinel Hociung, Petter Hohle, Jena Hwang, Takumi Ikeda, Radu Ion, Elena Irimia, Olájídé Ishola, Tomáš Jelínek, Anders Johannsen, Fredrik Jørgensen, Markus Juutinen, Hüner Kaşıkara, Andre Kaasen, Nadezhda Kabaeva, Sylvain Kahane, Hiroshi Kanayama, Jenna Kanerva, Boris Katz, Tolga Kayadelen, Jessica Kenney, Václava Kettnerová, Jesse Kirchner, Elena Klementieva, Arne Köhn, Kamil Kopacewicz, Natalia Kotsyba, Jolanta Kovalevskaite, Simon Krek, Sookyoung Kwak, Veronika Laippala, Lorenzo Lambertino, Lucia Lam, Tatiana Lando, Septina Dian Larasati, Alexei Lavrentiev, John Lee, Phương Lê Hông, Alessandro Lenci, Saran Lertpradit, Herman Leung, Cheuk Ying Li, Josie Li, Keying Li, KyungTae Lim, Maria Liovina, Yuan Li, Nikola Ljubešic, Olga Loginova, Olga Lyashevskaya, Teresa Lynn, Vivien Macketanz, Aibek Makazhanov, Michael Mandl, Christopher Manning, Ruli Manurung, Cătălina Mărănduc, David Marecek, Katrin Marheinecke, Héctor Martínez Alonso, André Martins, Jan Mašek, Yuji Matsumoto, Ryan McDonald, Sarah McGuinness, Gustavo Mendonça, Niko Miekka, Margarita Misirpashayeva, Anna Missilä, Cătălin Mititelu, Maria Mitrofan, Yusuke Miyao, Simonetta Montemagni, Amir More, Laura Moreno Romero, Keiko Sophie Mori, Tomohiko Morioka, Shinsuke Mori, Shigeki Moro, Bjartur Mortensen, Bohdan Moskalevskyi, Kadri Muischnek, Robert Munro, Yugo Murawaki, Kaili Müürisep, Pinkey Nainwani, Juan Ignacio Navarro Horñiacek, Anna Nedoluzhko, Gunta Nešpore-Berzkalne, Lương Nguyễn Thị, Huyền Nguyễn Thị Minh, Yoshihiro Nikaido, Vitaly Nikolaev, Rattima Nitisaroj, Hanna Nurmi, Stina Ojala, Atul Kr. Ojha, Adédayo Olúòkun, Mai Omura, Petya Osenova, Robert Östling, Lilja Øvrelid, Niko Partanen, Elena Pascual, Marco Passarotti, Agnieszka Patejuk, Guilherme Paulino-Passos, Angelika Peljak-Łapinska, Siyao Peng, Cenel-Augusto Perez, Guy Perrier, Daria Petrova, Slav Petrov, Jason Phelan, Jussi Piitulainen, Tommi A Pirinen, Emily Pitler, Barbara Plank, Thierry Poibeau, Larisa Ponomareva, Martin Popel, Lauma Pretkalniņa, Sophie Prévost, Prokopis Prokopidis, Adam Przepiórkowski, Tiina Puolakainen, Sampo Pyysalo, Peng Qi, Andriela Rääbis, Alexandre Rademaker, Loganathan Ramasamy, Taraka Rama, Carlos Ramisch, Vinit Ravishankar, Livy Real, Siva Reddy, Georg Rehm, Ivan Riabov, Michael Rießler, Erika Rimkute, Larissa Rinaldi, Laura Rituma, Luisa Rocha, Mykhailo Romanenko, Rudolf Rosa, Davide Rovati, Valentin Rosca, Olga Rudina, Jack Rueter, Shoval Sadde, Benoît Sagot, Shadi Saleh, Alessio Salomoni, Tanja Samardžic, Stephanie Samson, Manuela Sanguinetti, Dage Särg, Baiba Saul̄ıte, Yanin Sawanakunanon, Nathan Schneider, Sebastian Schuster, Djamé Seddah, Wolfgang Seeker, Mojgan Seraji, Mo Shen, Atsuko Shimada, Hiroyuki Shirasu, Muh Shohibussirri, Dmitry Sichinava, Aline Silveira, Natalia Silveira, Maria Simi, Radu Simionescu, Katalin Simkó, Mária Šimková, Kiril Simov, Aaron Smith, Isabela Soares-Bastos, Carolyn Spadine, Antonio Stella, Milan Straka, Jana Strnadová, Alane Suhr, Umut Sulubacak, Shingo Suzuki, Zsolt Szántó, Dima Taji, Yuta Takahashi, Fabio Tamburini, Takaaki Tanaka, Isabelle Tellier, Guillaume Thomas, Liisi Torga, Trond Trosterud, Anna Trukhina, Reut Tsarfaty, Francis Tyers, Sumire Uematsu, Zdenka Urešová, Larraitz Uria, Hans Uszkoreit, Andrius Utka, Sowmya Vajjala, Daniel van Niekerk, Gertjan van Noord, Viktor Varga, Eric Villemonte de la Clergerie, Veronika Vincze, Lars Wallin, Abigail Walsh, Jing Xian Wang, Jonathan North Washington, Maximilan Wendt, Seyi Williams, Mats Wirén, Christian Wittern, Tsegay Woldemariam, Tak-sum Wong, Alina Wróblewska, Mary Yako, Naoki Yamazaki, Chunxiao Yan, Koichi Yasuoka, Marat M. Yavrumyan, Zhuoran Yu, Zdenek Žabokrtský, Amir Zeldes, Manying Zhang, and Hanzhi Zhu. 2019. Universal Dependencies 2.5. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), Faculty of Mathematics and Physics, Charles University.

Citation

MLA
Qi, P., et al. “Stanza: A Python Natural Language Processing Toolkit for Many Human Languages”. arXiv, 2020, http://arxiv.org/abs/2003.07082v2.
APA
Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. arXiv. http://arxiv.org/abs/2003.07082v2
Chicago
Qi, P., Y. Zhang, Y. Zhang, J. Bolton, and C. D. Manning. 2020. “Stanza: A Python Natural Language Processing Toolkit for Many Human Languages”. arXiv. http://arxiv.org/abs/2003.07082v2.
Harvard
Qi, P. et al. (2020) “Stanza: A Python Natural Language Processing Toolkit for Many Human Languages”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2003.07082v2.
Vancouver
1. Qi P, Zhang Y, Zhang Y, Bolton J, Manning CD (2020) Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. arXiv

BibTeX

@article{qi2020stanza,
  title = {Stanza: A Python Natural Language Processing Toolkit for Many Human Languages},
  author = {Qi, Peng and Zhang, Yuhao and Zhang, Yuhui and Bolton, Jason and Manning, Christopher D.},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2003.07082v2},
  eprint = {2003.07082}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/