GottBERT: a pure German Language Model

Raphael ScheibleJohann FreiFabian ThomczykHenry HePatric TippmannJochen KnausVictor JaravineFrank KramerMartin Boeker

article2024EMNLP104 citations

Presents GottBERT, an open-source, single-language RoBERTa model trained purely on German web data that outperforms existing multilingual and monolingual baselines across multiple classification and named entity recognition tasks.

Listen

Organizations deploying natural language processing often face trade-offs between the substantial computational cost of massive prompt-based models and the need for high task accuracy in specific languages. Pre-trained, language-specific models offer a resource-efficient alternative to both large generative systems and multilingual models, which frequently dilute performance across many languages. The article addresses the gap in dedicated German language representations by presenting GottBERT, the first single-language German model based on the optimized RoBERTa architecture, and examines whether systematically filtering raw web text improves downstream model performance.

To construct and test these models, the researchers pre-trained base and large variants of GottBERT on the German portion of the OSCAR web crawl dataset (145 gigabytes) using dedicated Tensor Processing Units (TPUs) and a German-tailored vocabulary of 52,000 subwords. In parallel, they developed filtered versions (fGottBERT) using a 121-gigabyte cleaned corpus where encoding artifacts, non-German content, and low-quality text were removed using a syntax-based support vector machine. The evaluation benchmarked these models against existing monolingual German and leading multilingual architectures across six standard tasks: two named entity recognition benchmarks (CoNLL 2003 and GermEval 2014), three text classification datasets (coarse- and fine-grained GermEval 2018, and 10kGNAD), and natural language inference on XNLI.

The findings show that GottBERT base models delivered top-tier performance, outperforming all competing base models in four out of six downstream tasks. In named entity recognition and coarse tweet classification, GottBERT base models led their category with F1 scores of 87.59% on GermEval 2014, 86.14% on CoNLL 2003, and 78.65% on GermEval 2018. For large-scale architectures, GottBERT achieved competitive results (such as 83.31% accuracy on XNLI and over 90.2% on news classification), though competitors like GELECTRA and GBERT achieved slightly higher peak scores on several tasks. Unexpectedly, pre-training on the syntactically cleaned corpus did not provide a meaningful or consistent performance advantage over the uncleaned baseline.

These results indicate that dedicated, single-language models provide an efficient, high-performing foundation for enterprise German language tasks without requiring massive computational footprints. The lack of performance gains from strict text cleaning further implies that expensive, syntax-level pre-filtering of large web crawls may not be economically justified, as the models tolerate minor web noise and may even benefit from the broader linguistic variance in raw data. The authors have open-sourced all GottBERT models under the MIT license, enabling immediate, cost-effective integration into commercial and academic pipelines.

Decision-makers should consider using GottBERT base models for efficient production deployments in German document processing, entity extraction, and classification. For future pre-training initiatives, development teams should prioritize increasing dataset diversity (integrating sources like legal, news, and encyclopedia corpora) and exploring techniques like whole-word masking rather than investing heavily in basic syntax filtering. Readers should note that large model configurations were evaluated using conservative learning rates due to resource constraints, and performance may vary when encountering specialized regional dialects or non-standard syntax without dedicated fine-tuning.

arXiv: 2012.02110
Cover for GottBERT: a pure German Language Model

Abstract

Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimized version, RoBERTa. While initial research focused on English, single-language models can be advantageous compared to multi-lingual ones in terms of pre-training effort, over-all resource efficiency or downstream task performance. Despite the growing popularity of prompt-based LLMs, more compute-efficient BERT-like models remain highly relevant. In this work, we present the first German single-language RoBERTa model, GottBERT, pre-trained exclusively on the German portion of the OSCAR dataset. Additionally, we investigated the impact of filtering the OSCAR corpus. GottBERT was pre-trained using fairseq and standard hyperparameters. We evaluated its performance on two Named Entity Recognition (NER) tasks (Conll 2003 and GermEval 2014) and three text classification tasks (GermEval 2018 fine and coarse, and 10kGNAD) against existing German BERT models and two multi-lingual models. Performance was measured using the F1 score and accuracy. The GottBERT base and large models showed competitive performance, with GottBERT leading among the base models in 4 of 6 tasks. Contrary to our expectation, the applied filtering did not significantly affect the results. To support the German NLP research community, we are releasing the GottBERT models under the MIT license.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • Training Data
  • Filtering OSCAR
  • Pre-processing
  • Pre-training
  • Downstream Tasks
  • 4 Results
  • 5 Discussion
  • 6 Conclusion
  • Acknowledgments
  • Limitations
  • References
  • A Filtering OSCAR
  • B Model Properties
  • C Perplexity
  • D Parameters

Knowls

  1. Knowl 1 — GottBERT Architecture and Pre-training Setup

    model/method

    GottBERT is a monolingual German transformer language model based on the RoBERTa architecture and pre-trained using Fairseq on Google Cloud TPUs.

    The model is released in two primary architectures:

    1. GottBERT Base: 125,985,024 parameters, featuring a 52,009 token vocabulary. Pre-trained for 100,000 update steps with an effective batch size of 8,192 sequences. The training employs a 10,000-step linear warmup reaching a peak learning rate of 0.00040.0004, followed by a polynomial learning rate decay to zero. The unfiltered base model was trained on a 256-core TPUv3 pod for approximately 1.2 days.

    2. GottBERT Large: 357,145,600 parameters, sharing the 52,009 vocabulary. Pre-trained for 100,000 update steps with a batch size of 8,192 sequences, a 10,000-step warmup, and a peak learning rate of 0.000150.00015 with polynomial decay to zero. Pre-training was conducted on a 128-core TPUv4 pod taking approximately 5.7 days.

    Due to TPU memory architecture requiring fixed allocations, training data was ingested as a continuous single token stream without packing by individual document boundaries. Pre-training was executed in full 32-bit floating-point precision.

  2. Knowl 2 — Document Syntactic Ratios for German Text Quality Filtering

    equation

    To identify meaningless documents, spam, and poorly formatted text in German web crawls without expensive semantic parsing, four document-level syntactic ratio metrics are computed for a document DD consisting of a sequence of tokens (t0,t1,…,tn−1,tn)(t_0, t_1, \dots, t_{n-1}, t_n) with total token count nn and word token subset W(D)=(w0,w1,…,wm)W(D) = (w_0, w_1, \dots, w_m):

    1. Stopword Ratio (rsr_s): rs=∑i=0n∑ts∈S[ti=ts]nr_s = \frac{\sum_{i=0}^n \sum_{t_s \in S} [t_i = t_s]}{n} where SS is the German stopword list from NLTK and [⋅][\cdot] is the Iverson bracket indicating token equality.

    2. Punctuation Ratio (rpr_p): rp=∣{ti∣ti∉W(D)}∣nr_p = \frac{|\{ t_i \mid t_i \notin W(D) \}|}{n} measuring the proportion of non-word (punctuation or symbol) tokens.

    3. Unique Words Ratio (rur_u): ru=∣{wi∈W(D)∣∀wj∈W(D)∖{wi}:wi≠wj}∣nr_u = \frac{|\{ w_i \in W(D) \mid \forall w_j \in W(D) \setminus \{w_i\}: w_i \neq w_j \}|}{n} measuring the frequency of single-occurrence words.

    4. Upper Token Ratio (rupr_{up}): rup=∑wi∈W(D)c(wi)nr_{up} = \frac{\sum_{w_i \in W(D)} c(w_i)}{n} where the capitalization indicator is defined as: c(w)={1if the first letter of word w is uppercase0otherwisec(w) = \begin{cases} 1 & \text{if the first letter of word } w \text{ is uppercase} \\ 0 & \text{otherwise} \end{cases} Capitalization of nouns is a grammatical rule in German; deviations in rupr_{up} combined with extreme stopword or punctuation ratios indicate machine-generated text, lists, or ungrammatical informal noise.

  3. Knowl 3 — German OSCAR Corpus Multi-Stage Cleaning Pipeline

    model/method

    The raw German portion of the OSCAR dataset (145 GB, ~21.5 billion words across ~459 million line-separated documents) contains formatting noise, corrupted encodings, non-German text, and spam. A four-stage filtering pipeline cleans the data into a 121 GB corpus (~18.1 billion words across ~382 million documents):

    1. Encoding Correction: Erroneous German umlauts caused by UTF-8 text being decoded as legacy charsets (such as ISO-8859-1 or Windows-1252) are normalized. Lines containing unrecoverable Unicode replacement characters are discarded. Text normalization is performed via clean-text.
    2. Syntactic and PII Filtering: Regular expressions strip phone numbers, email addresses, URLs, and emojis. Documents containing fewer than 40 characters are discarded.
    3. Language and Art Detection: An n-gram language detection algorithm (whatlang-rs) filters out non-German passages and ASCII art.
    4. One-Class SVM Document Classification: A single-class Support Vector Machine (SVM) trained on the four document syntactic ratios (rs,rp,ru,rupr_s, r_p, r_u, r_{up}) separates clean documents from spam (such as word lists, raw code, and unspaced text). Trained unsupervised on 12,000 documents and evaluated on 1,750 annotated validation documents (15.5% dirty, 84.5% clean), the optimized SVM achieves a weighted F1F_1 score of 0.85780.8578 and a Matthews Correlation Coefficient (MCC) of 0.56630.5663.
  4. Knowl 4 — Language-Specific Subword Tokenization for German RoBERTa

    model/method

    Unlike BERT models that rely on WordPiece tokenizers with language-specific pre-tokenization tools (such as Moses), GottBERT uses GPT-2 Byte Pair Encoding (BPE) operating directly on raw text bytes.

    To avoid token representation inefficiencies caused by English-centric vocabularies, a dedicated German vocabulary of 52,009 subword units was constructed from a randomly sampled 40 GB subset of the German OSCAR corpus. This language-specific subword vocabulary achieves a 40% reduction in binary data size during tokenization compared to the standard English GPT-2 tokenizer while improving representation quality for German compounds and morphology.

  5. Knowl 5 — Benchmark Results Across German Downstream Tasks

    data/table

    Downstream task performance comparing base and large variants of GottBERT (unfiltered OSCAR), fGottBERT^\text{f}\text{GottBERT} (filtered OSCAR), and final epoch checkpoints (denoted with †\dagger) against existing German single-language models (GBERT, GELECTRA, dbmdzBERT, GermanBERT) and multilingual models (mBERT, XLM-RoBERTa). Metrics reported are accuracy (%) for XNLI and macro F1F_1 score (%) for NER and classification tasks, selected as the best score out of 24 grid-search runs per model on the validation set.

    Model XNLI GermEval 2014 CoNLL 03 GermEval 2018 10kGNAD
    coarse fine
    GottBERT_base 80.82 87.55 85.93 78.17 53.30 89.64
    GottBERT^_base 81.04 87.48 85.61 78.18 53.92 90.27
    ^fGottBERT_base 80.56 87.57 86.14 78.65 52.82 89.79
    ^fGottBERT^_base 80.74 87.59 85.66 78.08 52.39 89.92
    GELECTRA_base 81.70 86.91 85.37 77.26 50.07 89.02
    GBERT_base 80.06 87.24 85.16 77.37 51.51 90.30
    dbmdzBERT 68.12 86.82 85.15 77.46 52.07 90.34
    GermanBERT 78.16 86.53 83.87 74.81 47.78 90.18
    XLM-R_base 79.76 86.14 84.46 77.13 50.54 89.81
    mBERT 77.03 86.67 83.18 73.54 48.32 88.90
    GottBERT_large 82.46 88.20 86.78 79.40 54.61 90.24
    ^fGottBERT_large 83.31 88.13 86.30 79.32 54.70 90.31
    ^fGottBERT^_large 82.79 88.27 86.28 78.96 54.72 90.17
    GELECTRA_large 86.33 88.72 86.78 81.28 56.17 90.97
    GBERT_large 84.21 88.72 87.19 80.84 57.37 90.74
    XLM-R_large 84.07 88.83 86.54 79.05 55.06 90.17

    The evaluation shows that among base architectures, GottBERT models achieve the highest score in 4 out of 6 tasks (GermEval 2014, CoNLL 03, GermEval 2018 coarse, GermEval 2018 fine). Among large architectures, GELECTRA and GBERT lead across NLI and classification tasks.

  6. Knowl 6 — Downstream Fine-Tuning Hyperparameter Optimization Protocol

    experimental setup

    Evaluation on downstream tasks is performed by executing a grid search over 24 runs per model and task combination using the HuggingFace Transformers library.

    The search space comprises:

    • Learning Rates: {5×10−5,2×10−5,1×10−5,7×10−6,5×10−6,1×10−6}\{5\times 10^{-5}, 2\times 10^{-5}, 1\times 10^{-5}, 7\times 10^{-6}, 5\times 10^{-6}, 1\times 10^{-6}\}
    • Batch Sizes: {16,32,48,64}\{16, 32, 48, 64\}
    • Epoch Limits: Up to 30 epochs for Named Entity Recognition (GermEval 2014, CoNLL 2003) and Text Classification (GermEval 2018 coarse/fine, 10kGNAD); up to 10 epochs for Natural Language Inference (XNLI).

    Model selection is determined by the highest validation metric (F1F_1 for NER/CLS, accuracy for NLI) across the 24 runs, and reported results correspond to test-set evaluations of that selected checkpoint. Downstream evaluation consumed 1,549 hours and 29 minutes (~64.6 GPU days) on Nvidia Titan RTX and Nvidia A40 GPUs.

  7. Knowl 7 — Impact of Heuristic Web Corpus Filtering on Downstream Performance

    empirical result

    Comparing GottBERT models pre-trained on the raw 145 GB German OSCAR corpus versus fGottBERT^\text{f}\text{GottBERT} models pre-trained on the 121 GB filtered corpus demonstrates that aggressive syntactic and heuristic data cleaning does not provide a consistent downstream performance advantage.

    Although filtered models underwent additional training epochs over the smaller dataset (15.87 epochs for filtered base vs. 13.07 epochs for unfiltered base), performance differences were marginal across downstream tasks:

    • In base models, fGottBERT^\text{f}\text{GottBERT} achieved slight gains on GermEval 2014 (87.59%87.59\% vs. 87.55%87.55\%) and GermEval 2018 coarse (78.65%78.65\% vs. 78.18%78.18\%).
    • Unfiltered GottBERT base performed better on GermEval 2018 fine (53.92%53.92\% vs. 52.82%52.82\%) and 10kGNAD (90.27%90.27\% vs. 89.92%89.92\%).
    • In large models, fGottBERTlarge^\text{f}\text{GottBERT}_{\text{large}} showed a slight gain on XNLI (83.31%83.31\% vs. 82.46%82.46\%), while unfiltered GottBERT held advantages on CoNLL 03 (86.78%86.78\% vs. 86.30%86.30\%) and GermEval 2018 coarse (79.40%79.40\% vs. 79.32%79.32\%).

    Heuristic filtering may inadvertently remove text diversity and lexical variance, which counteracts the theoretical benefits of eliminating noise.

  8. Knowl 8 — Pre-training Perplexity Dynamics and TPU Optimization Constraints

    empirical result

    Tracking the pre-training perplexity of GottBERT over optimization cycles on TPU reveals two key dynamics:

    1. Convergence Behavior: Both base and large architectures display an initial training plateau followed by flat convergence after approximately 40,000 update steps. Large models experience longer plateau durations and transient upward spikes in batch-level perplexity during pre-training before stabilizing.

    2. TPU Compute and Implementation Constraints: Because TPUs enforce static memory layouts without dynamic allocation, text inputs must be supplied as a continuous 1D token stream rather than boundary-aware document segments. Furthermore, 32-bit floating point precision was required throughout training, increasing hardware memory pressure and limiting the feasibility of exploring broader pre-training learning rate schedules for the large model.

  9. Knowl 9 — Architectural and Methodological Limitations of GottBERT

    limitation

    The methodology and evaluation of GottBERT have five identified limitations:

    1. Heuristic Syntactic Cleaning: Data filtering relies strictly on shallow syntactic features and regular expressions; semantic noise, privacy-sensitive personal information, and subtle factual errors are not detected or removed.
    2. Corpus Generalization and Dialectal Coverage: Training is restricted to the first version of the OSCAR crawl, potentially limiting robustness against non-standard German dialects, Austrian/Swiss regional variations, and evolving online language.
    3. Absence of Whole-Word Masking (WWM): Pre-training employs standard subword-level masking rather than whole-word masking, which has been shown to boost downstream contextual representation quality in successor models like GBERT.
    4. Restricted Pre-training Hyperparameter Tuning: Due to substantial compute requirements (~5.7 days on 128-core TPUv4 per large run), large models were trained using a single conservative peak learning rate (0.000150.00015) without learning rate search.
    5. Single-Annotator Quality Validation: The One-Class SVM quality filter was validated against a 1,750-document set labeled by a single annotator, introducing potential subjective bias in clean vs. dirty document categorization.

Coverage note — No substantial contributed material was omitted. All key contributions—including model architecture, tokenizer setup, corpus cleaning equations and pipeline, downstream evaluation protocol, tabular benchmark results, pre-training loss dynamics, and stated limitations—are fully covered.

References

  1. 1.Darina Benikova, Chris Biemann, Max Kisselew, and Sebastian Padó. 2014. GermEval 2014 Named Entity Recognition Shared Task: Companion Paper. Proceedings of the KONVENS GermEval Shared Task on Named Entity Recognition, pages 104–112.
  2. 2.Keno K. Bressem, Jens-Michalis Papaioannou, Paul Grundmann, Florian Borchert, Lisa C. Adams, Leonhard Liu, Felix Busch, Lina Xu, Jan P. Loyen, Stefan M. Niehues, Moritz Augustin, Lennart Grosser, Marcus R. Makowski, Hugo J.W.L. Aerts, and Alexander Löser. 2024. medbert.de: A comprehensive german bert model for the medical domain. Expert Systems with Applications, 237:121598.
  3. 3.William B. Cavnar and John M. Trenkle. 1994. N-Gram-Based Text Categorization. In In Proceedings of SDAIR-94, 3rd Annual Symposium on Document Analysis and Information Retrieval, pages 161–175.
  4. 4.Branden Chan, Stefan Schweter, and Timo Möller. 2020. German’s next language model. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6788–6796, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  5. 5.Davide Chicco and Giuseppe Jurman. 2020. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21(1):6.
  6. 6.Davide Chicco, Matthijs J. Warrens, and Giuseppe Jurman. 2021. The Matthews Correlation Coefficient (MCC) is More Informative Than Cohen’s Kappa and Brier Score in Binary Classification Assessment. IEEE Access, 9:78368–78381. Conference Name: IEEE Access.
  7. 7.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. Electra: Pre-training text encoders as discriminators rather than generators. Preprint, arXiv:2003.10555.
  8. 8.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Unsupervised cross-lingual representation learning at scale. CoRR, abs/1911.02116.
  9. 9.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  10. 10.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  11. 11.David Crystal and Honorary Professor of Linguistics David Crystal. 2003. The Cambridge Encyclopedia of the English Language. Cambridge University Press.
  12. 12.Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019. BERTje: A Dutch BERT Model. arXiv:1912.09582 [cs]. ArXiv: 1912.09582.
  13. 13.Pieter Delobelle, Thomas Winters, and Bettina Berendt. 2020. RobBERT: a Dutch RoBERTa-based Language Model. arXiv:2001.06286 [cs]. ArXiv: 2001.06286.
  14. 14.Jacob Devlin. 2018. Multilingual BERT Readme Document.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020. Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping. arXiv:2002.06305 [cs]. ArXiv: 2002.06305.
  17. 17.Johann Frei, Ludwig Frei-Stuber, and Frank Kramer. 2022. Gernermed++: Transfer learning in german medical nlp. Preprint, arXiv:2206.14504.
  18. 18.Johann Frei and Frank Kramer. 2023. Annotated dataset creation through large language models for non-english medical nlp. Journal of Biomedical Informatics, 145:104478.
  19. 19.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. Preprint, arXiv:2310.06825.
  20. 20.Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. 2023. TPU v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. Preprint, arxiv:2304.01433 [cs].
  21. 21.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondˇrej Bojar, Alexandra Constantin, and Evan Herbst. 2007. Moses: Open Source Toolkit for Statistical Machine Translation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics Companion Volume Proceedings of the Demo and Poster Sessions, pages 177–180, Prague, Czech Republic. Association for Computational Linguistics.
  22. 22.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  23. 23.Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, and Didier Schwab. 2020. FlauBERT: Unsupervised Language Model Pre-training for French. arXiv:1912.05372 [cs]. ArXiv: 1912.05372.
  24. 24.Manuel Lentzen, Sumit Madan, Vanessa Lage-Rupprecht, Lisa Kühnel, Juliane Fluck, Marc Jacobs, Mirja Mittermaier, Martin Witzenrath, Peter Brunecker, Martin Hofmann-Apitius, Joachim Weber, and Holger Fröhlich. 2022. Critical assessment of transformer-based AI models for German clinical notes. JAMIA Open, 5(4):ooac087.
  25. 25.Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, Zihao Wu, Lin Zhao, Dajiang Zhu, Xiang Li, Ning Qiang, Dingang Shen, Tianming Liu, and Bao Ge. 2023. Summary of chatgpt-related research and perspective towards the future of large language models. Meta-Radiology, 1(2):100017.
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs]. ArXiv: 1907.11692.
  27. 27.Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020. CamemBERT: a Tasty French Language Model. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7203–7219, Online. Association for Computational Linguistics.
  28. 28.Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019. Facebook FAIR’s WMT19 News Translation Task Submission. In Proceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), pages 314–319, Florence, Italy. Association for Computational Linguistics.
  29. 29.Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020. A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1703–1714, Online. Association for Computational Linguistics.
  30. 30.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019. fairseq: A Fast, Extensible Toolkit for Sequence Modeling. arXiv:1904.01038 [cs]. ArXiv: 1904.01038.
  31. 31.Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018. Scaling Neural Machine Translation. arXiv:1806.00187 [cs]. ArXiv: 1806.00187.
  32. 32.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9.
  33. 33.Julian Risch, Eva Krebs, Alexander Löser, Alexander Riese, and Ralf Krestel. 2018. Fine-Grained Classification of Offensive Language. In Proceedings of GermEval 2018 (co-located with KONVENS), pages 38–44.
  34. 34.Dietmar Schabus, Marcin Skowron, and Martin Trapp. 2017. One Million Posts: A Data Set of German Online Discussions. In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), pages 1241–1244, Tokyo, Japan.
  35. 35.Raphael Scheible, Fabian Thomczyk, Patric Tippmann, Victor Jaravine, and Martin Boeker. 2020. Gottbert: a pure german language model. Preprint, arXiv:2012.02110.
  36. 36.Moritz Scherrmann. 2023. German finbert: A german pre-trained language model. Preprint, arXiv:2311.08793.
  37. 37.B Schölkopf, R Williamson, AJ Smola, and J Shawe-Taylor. 1999. Single-class support vector machines. In Dagstuhl-Seminar 99121: Unsupervised Learning, pages 19–20. Schloss Dagstuhl, Leibniz-Zentrum für Informatik.
  38. 38.Mike Schuster and Kaisuke Nakajima. 2012. Japanese and Korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152. ISSN: 2379-190X.
  39. 39.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural Machine Translation of Rare Words with Subword Units. arXiv:1508.07909 [cs]. ArXiv: 1508.07909.
  40. 40.Daniel Solling. 2009. Små bokstäver ökade avståndet till tyskarna. Library Catalog: spraktidningen.se.
  41. 41.Erik F. Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003 - Volume 4, CONLL ’03, pages 142–147, USA. Association for Computational Linguistics. Event-place: Edmonton, Canada.
  42. 42.Cagri Toraman, Eyup Halit Yilmaz, ¸Sahinuç Furkan, and Oguzhan Ozcelik. 2023. Impact of tokenization on language models: An analysis for turkish. ACM Trans. Asian Low-Resour. Lang. Inf. Process., 22(4).
  43. 43.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. Preprint, arXiv:2302.13971.
  44. 44.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems 30, pages 5998–6008. Curran Associates, Inc.
  46. 46.Antti Virtanen, Jenna Kanerva, Rami Ilo, Jouni Luoma, Juhani Luotolahti, Tapio Salakoski, Filip Ginter, and Sampo Pyysalo. 2019. Multilingual is not enough: BERT for Finnish. arXiv:1912.07076 [cs]. ArXiv: 1912.07076.
  47. 47.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  48. 48.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019. Huggingface’s transformers: State-of-the-art natural language processing. CoRR, abs/1910.03771.
  49. 49.Haoran Xu, Benjamin Van Durme, and Kenton Murray. 2021. Bert, mbert, or bibert? a study on contextualized embeddings for neural machine translation. Preprint, arXiv:2109.04588.
  50. 50.Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020. Large Batch Optimization for Deep Learning: Training BERT in 76 minutes. arXiv:1904.00962 [cs, stat]. ArXiv: 1904.00962 version: 5.

Citation

MLA
Scheible, R., et al. “GottBERT: A Pure German Language Model”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 21237–50, https://doi.org/10.18653/v1/2024.emnlp-main.1183.
APA
Scheible, R., Frei, J., Thomczyk, F., He, H., Tippmann, P., Knaus, J., Jaravine, V., Kramer, F., & Boeker, M. (2024). GottBERT: a pure German Language Model. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21237–21250. https://doi.org/10.18653/v1/2024.emnlp-main.1183
Chicago
Scheible, R., J. Frei, F. Thomczyk, et al. 2024. “GottBERT: A Pure German Language Model”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21237–50. https://doi.org/10.18653/v1/2024.emnlp-main.1183.
Harvard
Scheible, R. et al. (2024) “GottBERT: a pure German Language Model”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 21237–21250. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.1183.
Vancouver
1. Scheible R, Frei J, Thomczyk F, He H, Tippmann P, Knaus J, Jaravine V, Kramer F, Boeker M (2024) GottBERT: a pure German Language Model. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 21237–21250

BibTeX

@inproceedings{Scheible_2024, title={GottBERT: a pure German Language Model}, url={http://dx.doi.org/10.18653/v1/2024.emnlp-main.1183}, DOI={10.18653/v1/2024.emnlp-main.1183}, booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Scheible, Raphael and Frei, Johann and Thomczyk, Fabian and He, Henry and Tippmann, Patric and Knaus, Jochen and Jaravine, Victor and Kramer, Frank and Boeker, Martin}, year={2024}, pages={21237–21250} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/