FinGPT: Large Generative Models for a Small Language

Risto LuukkonenVille KomulainenJouni LuomaAnni EskelinenJenna KanervaHanna-Mari KupariFilip GinterVeronika LaippalaNiklas MuennighoffAleksandra Piktus

article2023EMNLP74 citations

Presents open-source Finnish generative language models scaling up to 176 billion parameters alongside a localized evaluation benchmark, offering a practical framework for developing large-scale foundation models for lower-resourced languages.

Listen

Most modern large language models require hundreds of billions of words of training text and disproportionately focus on English, leaving smaller languages underserved. Finnish is natively spoken by fewer than six million people and accounts for less than one percent of common online text repositories, making the development of capable generative artificial intelligence challenging under severe data constraints.

The article demonstrates effective strategies for developing advanced generative language models for data-constrained languages by training seven monolingual Finnish models from scratch and adapting an existing massive multilingual model for Finnish capabilities.

To establish credibility and broad linguistic coverage, the authors compiled a 207-billion-character dataset combining web crawls, national news archives, discussion forums, social media, and electronic books. After filtering for quality, perplexity, and toxicity, the preprocessed corpus yielded 38 billion tokens. Using the LUMI supercomputer, the authors trained seven monolingual models ranging from 186 million to 13 billion parameters over roughly eight passes (300 billion tokens). Additionally, they performed continued training on the 176-billion-parameter multilingual BLOOM model using a mix of its original training data and Finnish text, resulting in an extended model named BLUUMI. Model evaluation was conducted using FIN-bench, a newly adapted benchmark comprising 3,919 Finnish test examples across multiple reasoning and language tasks.

The analysis produced several key findings. First, the 176-billion-parameter BLUUMI model achieved the highest overall benchmark performance, outperforming the best prior Finnish model by over 20 percentage points and the original multilingual model by 12 to 18 percentage points on Finnish tasks, without degrading its English proficiency. Second, the monolingual 8-billion-parameter model delivered the strongest performance among the models trained from scratch, outperforming previous Finnish models of comparable size by over 10 percentage points. Third, performance declined between the 8-billion and 13-billion-parameter monolingual models, indicating that roughly 10 billion parameters represents an upper limit when training from scratch on limited text before overfitting reduces capability. Fourth, toxicity filtering of training data reduced the rate of toxic text generation by more than half compared to unfiltered models, although models still generated toxic output roughly 2% of the time. Finally, all evaluated base models performed poorly on basic alignment tests measuring helpfulness, honesty, and harmlessness.

These findings prove that high-performing foundation models can be successfully developed for lesser-resourced languages through rigorous data curation and extended pretraining of existing multilingual models. The results indicate that adapting existing large multilingual models is more effective than training large monolingual models from scratch when language data is limited. However, because these systems retain occupational gender biases and fail baseline safety and helpfulness evaluations, deploying them in unaligned forms presents significant compliance, reputational, and operational risks.

Organizations should treat these released models strictly as research foundations rather than production-ready systems. They should not be deployed in user-facing settings or high-stakes decision workflows, such as hiring, without substantial secondary intervention. Stakeholders pursuing language model implementations in data-constrained languages should prioritize continued pretraining of existing multilingual models over building large standalone models from scratch, and should incorporate explicit instruction tuning and safety alignment before operational release.

The findings are subject to specific limitations. The available volume of unique Finnish text forced repetitive training over multiple epochs, which contributed to performance degradation in the largest monolingual model. Furthermore, full replication of the training corpus is restricted because high-quality electronic book and newspaper data from the National Library of Finland cannot be redistributed due to copyright. Consequently, readers can have high confidence in the benchmarked technical capabilities but must remain cautious regarding the models' readiness for direct deployment.

arXiv: 2311.05640
Cover for FinGPT: Large Generative Models for a Small Language

Abstract

Large language models (LLMs) excel in many tasks in NLP and beyond, but most open models have very limited coverage of smaller languages and LLM work tends to focus on languages where nearly unlimited data is available for pretraining. In this work, we study the challenges of creating LLMs for Finnish, a language spoken by less than 0.1% of the world population. We compile an extensive dataset of Finnish combining web crawls, news, social media and eBooks. We pursue two approaches to pretrain models: 1) we train seven monolingual models from scratch (186M to 13B parameters) dubbed FinGPT, 2) we continue the pretraining of the multilingual BLOOM model on a mix of its original training data and Finnish, resulting in a 176 billion parameter model we call BLUUMI. For model evaluation, we introduce FIN-bench, a version of BIG-bench with Finnish tasks. We also assess other model qualities such as toxicity and bias. Our models and tools are openly available at https://turkunlp.org/gpt3-finnish.

Table of Contents

  • 1 Introduction
  • 2 Models
  • 3 Data
  • 3.1 Data sources
  • 3.2 Preprocessing
  • 3.3 Data statistics
  • 3.4 Register analysis
  • 4 Pretraining
  • 5 Evaluation
  • 5.1 FIN-bench dataset
  • 5.2 Few-shot results
  • 5.3 Alignment
  • 5.4 Bias
  • 5.5 Toxicity
  • 6 Discussion and conclusions
  • Acknowledgments
  • Limitations
  • References
  • A Timespan covered by Finnish datasets
  • B Comparison of mC4-Fi and CC-Fi datasets
  • C Full gender bias results on 13B model
  • D Toxicity scores
  • E Data distribution by source before and after weighting
  • F Full FIN-bench evaluation results
  • G FIN-bench examples

Knowls

  1. Knowl 1 — FinGPT Monolingual Model Family Architecture and Pretraining Configurations

    model/method

    The FinGPT family comprises seven decoder-only autoregressive language models trained from scratch exclusively on Finnish text for 300 billion tokens (spanning approximately 8 epochs over a 38-billion-token corpus). The model architectures follow the parameter, layer, and attention-head configurations of GPT-3 while adopting architectural features from BLOOM, including ALiBi (Attention with Linear Biases) positional embeddings and pre-layer normalization.

    Training was executed using Megatron-DeepSpeed adapted for AMD Instinct MI250X GPUs via ROCm (utilizing rocBLAS, MIOpen, and RCCL over the Slingshot interconnect with the AWS OFI plugin) on the LUMI supercomputer across up to 192 nodes (1,536 GCDs).

    Model Layers Dim Heads Parameters Batch Size (samples) Batch Size (tokens) Peak Learning Rate
    Small 12 768 12 186M 256 524,288 6.0×10−46.0 \times 10^{-4}
    Medium 24 1024 16 437M 256 524,288 3.0×10−43.0 \times 10^{-4}
    Large 24 1536 16 881M 256 524,288 2.5×10−42.5 \times 10^{-4}
    XL 24 2064 24 1.5B 512 1,048,576 2.0×10−42.0 \times 10^{-4}
    3B 32 2560 32 2.8B 512 1,048,576 1.6×10−41.6 \times 10^{-4}
    8B 32 4096 32 7.5B 1024 2,097,152 1.2×10−41.2 \times 10^{-4}
    13B 40 5120 40 13.3B 1024 2,097,152 1.0×10−41.0 \times 10^{-4}
  2. Knowl 2 — BLUUMI: Continued Pretraining of Multilingual BLOOM on Finnish

    model/method

    BLUUMI is a 176-billion-parameter multilingual model produced by continuing the pretraining of the base BLOOM-176B checkpoint on a blended corpus consisting of its original ROOTS pretraining dataset and cleaned Finnish text. The architecture retains the 70 transformer layers, hidden dimension of 14,336, 112 attention heads, and ALiBi positional embeddings of BLOOM.

    Pretraining hyperparameters were preserved from the original BLOOM recipe, using a batch size of 2,048 sequences (4,194,304 tokens per step) and a maximum learning rate of 6.0×10−56.0 \times 10^{-5}. Continued pretraining addresses the undertraining of the original BLOOM checkpoint while expanding its capability to Finnish without compromising its existing capabilities in other languages.

  3. Knowl 3 — Finnish Pretraining Corpus Assembly and Preprocessing Pipeline

    model/method

    A broad Finnish pretraining dataset totaling 207.0 billion characters (38 billion tokens after weighting and tokenization) was constructed from 11 text sources:

    • Web crawls: Finnish Internet Parsebank (35.0B chars, 1.5×\times weight), mC4 Finnish subset (46.3B chars, 1.0×\times weight), CC-Fi extracted from 2013-2022 Common Crawls with Trafilatura and MurmurHash deduplication (79.6B chars, 1.0×\times weight).
    • High-quality curated text: Finnish Wikipedia (0.8B chars, 3.0×\times weight), Projekti Lönnrot out-of-copyright literature (0.8B chars, 3.0×\times weight), Yle news archives (1.6B chars, 2.0×\times weight), STT Finnish News Agency archives (2.2B chars, 2.0×\times weight).
    • Books and archived print: National Library ePub eBooks (13.5B chars, 1.0×\times weight), National Library Lehdet newspaper/domain crawls (5.8B chars, 1.0×\times weight).
    • Social discussion: Suomi24 discussion forums (20.6B chars, 1.0×\times weight), Reddit r/Suomi posts and comments (0.7B chars, 1.0×\times weight).

    Preprocessing consists of five sequential filtering stages:

    1. N-gram deduplication: Onion marks duplicate lines if ≥50%\ge 50\% of N-grams occurred previously; whole documents are discarded if ≥50%\ge 50\% of lines are duplicates after trimming edge duplicates.
    2. Heuristic rule filtering: Removes texts with abnormal punctuation-to-alpha or digit-to-alpha ratios, low type-token ratios, or high non-Finnish character fractions (retaining >95%>95\% of data).
    3. Perplexity filtering: A 5-gram KenLM language model trained on FinBERT clean data discards non-standard prose lines with perplexity >100,000>100,000 (removing 15--20% of web crawl text and 2--5% of social media text; omitted for news, literature, and Wikipedia).
    4. Toxicity filtering: A FinBERT classifier fine-tuned on Finnish-translated Jigsaw toxicity data discards toxic documents (removing 1--5% of text from most sources and 23% from CC-Fi; not applied to news, literature, or Wikipedia).
    5. PII masking: Rule-based regex scripts mask email addresses and telephone numbers (affecting ≈0.2%\approx 0.2\% of characters).

    Tokenization uses a byte-level BPE tokenizer without Unicode normalization, featuring a custom vocabulary of 131,072 tokens to handle the rich morphology and compound structure of Finnish.

  4. Knowl 4 — FIN-bench Multi-Task Benchmark Dataset

    experimental setup

    FIN-bench is a standardized Finnish few-shot evaluation suite consisting of 3,919 examples across 11 diverse natural language understanding and reasoning tasks. Most tasks originate from BIG-bench via machine translation followed by human manual correction and cultural adaptation, augmented with new and domain-specific tasks:

    • Analogy (130 examples): Natural language analogy questions (AA is to BB as CC is to DD) derived from Finnish semantic resources.
    • Arithmetic (1,923 examples): One- to five-digit addition, subtraction, multiplication, and division problems generated via translated BIG-bench templates.
    • Cause and Effect (153 examples): Distinguishing causality direction and identifying causes and effects across one- and two-sentence variants.
    • Emotions (160 examples): Emotion classification on curated sentences of ≥5\ge 5 words from the Finnish XED dataset.
    • Empirical Judgments (99 examples): Categorizing statements as causal, correlative, or neutral non-relational.
    • General Knowledge (70 examples): Common sense factual questions translated with imperial units converted to metric.
    • Intent Recognition (693 examples): Identifying intent from user utterances, adapted from task-oriented dialogue data.
    • Misconceptions (134 examples): Discriminating true facts from popular misconceptions, with U.S.-centric questions filtered out.
    • Paraphrase (200 examples): Discriminating true paraphrases from topical semantic similarity using Finnish Paraphrase Corpus examples.
    • Sentence Ambiguity (60 examples): Evaluating truth claims in sentences with intentionally ambiguous syntax or semantics.
    • Similarities Abstraction (76 examples): Multiple-choice human-like abstract association identification between paired concepts.
  5. Knowl 5 — Few-Shot Capabilities and Monolingual Scaling Limit on FIN-bench

    empirical result

    Across zero- to three-shot evaluation on FIN-bench, the monolingual FinGPT models and BLUUMI outperform all prior Finnish open models across all parameter scales:

    • The prior state-of-the-art Finnish model (Hatanpää XL, ≈1.5\approx 1.5B parameters) achieved a maximum aggregate 3-shot accuracy of 38.5%.
    • FinGPT 8B achieves ≈50.3%\approx 50.3\% accuracy in 3-shot settings (an improvement of >10>10 percentage points over prior models), while BLUUMI (176B) reaches ≈59.2%\approx 59.2\% aggregate accuracy (>20>20 percentage points higher than prior Finnish models).
    • Scaling inversion occurs between FinGPT 8B (7.5B parameters) and FinGPT 13B (13.3B parameters), with the 13B model demonstrating a consistent drop across 0-shot to 3-shot accuracy (3-shot accuracy dropping from ≈50.3%\approx 50.3\% to ≈48.9%\approx 48.9\%).

    Because training up to 300 billion tokens on a 38-billion-token corpus requires ≈8\approx 8 epochs, larger parameter models suffer from multi-epoch overfitting on non-repeated text. This establishes approximately 10B parameters as an empirical upper limit for training monolingual autoregressive models from scratch under Finnish-scale pretraining data availability.

  6. Knowl 6 — Finnish Capability Expansion and English Knowledge Retention in BLUUMI

    empirical result

    Continued pretraining of BLOOM-176B on a joint corpus of ROOTS and Finnish text yields significant Finnish capability gains while preserving base English performance:

    • On FIN-bench, BLUUMI improves over BLOOM by 12 to 18 percentage points across all evaluation conditions: 0-shot accuracy increases from ≈36.1%\approx 36.1\% to ≈47.9%\approx 47.9\%, 1-shot from ≈38.2%\approx 38.2\% to ≈56.4%\approx 56.4\%, 2-shot from ≈40.5%\approx 40.5\% to ≈57.9%\approx 57.9\%, and 3-shot from ≈41.3%\approx 41.3\% to ≈59.2%\approx 59.2\%.
    • On 27 standard English language tasks from the EleutherAI LM Evaluation Harness (including ARC Easy/Challenge, BoolQ, COPA, HellaSwag, LAMBADA, MathQA, MRPC, OpenBookQA, PIQA, PubMedQA, QNLI, QQP, RACE, RTE, SciQ, SST, TriviaQA, WebQS, Winogrande, WNLI, and WSC), a two-sided tt-test shows no statistically significant performance difference between BLOOM and BLUUMI.

    This demonstrates that continued pretraining on a joint mixture of prior multilingual data and a new target language successfully incorporates the new language without catastrophic forgetting.

  7. Knowl 7 — Unprompted Toxicity Rates Across Pretraining Toxicity Filtering Regimes

    empirical result

    Evaluating unprompted generation toxicity (prompting language models exclusively with the end-of-sequence token and classifying outputs with a FinBERT toxicity detector) demonstrates that pretraining toxicity filtering substantially reduces toxic output generation:

    • Monolingual FinGPT models trained on data filtered for toxicity generate toxic output in 1.66%1.66\% to 2.37%2.37\% of unprompted generations.
    • Prior Finnish models trained without pretraining toxicity filtering (Hatanpää Small and XL) generate toxic output in 5.38%5.38\% and 5.24%5.24\% of generations, over double the rate of FinGPT.
    Model Identity Attack Insult Obscene Severe Toxicity Threat Total Toxicity
    Hatanpää/small 0.149% 1.471% 2.132% 0.070% 0.026% 5.377%
    Hatanpää/xl 0.185% 1.344% 2.055% 0.109% 0.015% 5.241%
    TurkuNLP/small 0.039% 0.208% 0.435% 0.004% 0.008% 1.658%
    TurkuNLP/medium 0.048% 0.248% 0.410% 0.002% 0.011% 1.896%
    TurkuNLP/large 0.039% 0.280% 0.490% 0.001% 0.011% 1.981%
    TurkuNLP/xl 0.061% 0.272% 0.546% 0.002% 0.011% 2.211%
    TurkuNLP/3B 0.069% 0.343% 0.618% 0.004% 0.021% 2.290%
    TurkuNLP/8B 0.058% 0.304% 0.645% 0.012% 0.021% 2.317%
    TurkuNLP/13B 0.065% 0.309% 0.637% 0.005% 0.016% 2.374%
  8. Knowl 8 — Gender-Occupation Stereotyping in FinGPT vs. Labor Statistics

    empirical result

    When prompted with Finnish prefix templates of the form "The name of the [occupation holder] was" ("[Ammatti]n nimi oli"), the FinGPT 13B model predicts male and female names that broadly reflect real-world labor distributions while amplifying gender skews for multiple professions relative to Statistics Finland 2020 employment statistics:

    • Heavily skewed professions match official data directions: practical nurse is 88.8%88.8\% female in national statistics and 100.0%100.0\% female in model predictions; building electrician is 97.5%97.5\% male in statistics and 100.0%100.0\% male in model predictions; house builder is 94.2%94.2\% male in statistics and 99.4%99.4\% male in model predictions; nanny is 95.0%95.0\% female in statistics and 91.6%91.6\% female in model predictions.
    • Male bias amplification occurs across several moderately balanced or female-majority occupations: seller (65.3%65.3\% female in statistics vs. 78.1%78.1\% male in model predictions), office cleaner (83.0%83.0\% female in statistics vs. 68.2%68.2\% male in model predictions), senior physician (60.3%60.3\% female in statistics vs. 90.7%90.7\% male in model predictions), and bank worker (77.8%77.8\% female in statistics vs. 84.3%84.3\% male in model predictions).

    These results confirm that the pretraining corpus imprints occupational gender associations onto model generations, necessitating intervention before applying raw base models to decision-support applications.

  9. Knowl 9 — Alignment Failure of Unaligned Base Finnish LLMs on HHH Evaluation

    empirical result

    Evaluation of base pretrained Finnish models and BLUUMI on the translated Finnish BIG-bench HHH (Harmlessness, Honesty, Helpfulness, and Other) benchmark reveals poor alignment across all model sizes:

    • All models score near or below the 50% random choice baseline on harmlessness, honesty, and helpfulness, with helpfulness yielding the lowest accuracy (measuring between 20% and 35%).
    • Only the "other" category consistently exceeds the 50% random baseline (scoring between 55% and 60%).
    • Performance differences between model scales remain within one standard deviation and are not statistically significant.

    This indicates that scaling pretraining parameters and token counts alone does not induce helpful or harmless conversational behavior in base generative language models in the absence of instruction tuning or alignment finetuning.

  10. Knowl 10 — Data Availability, Epoch Limits, and Reproducibility Constraints for Small-Language LLMs

    limitation

    Pretraining large generative models for smaller languages such as Finnish faces fundamental data availability and replication constraints:

    1. Total clean text in Finnish is limited to 38 billion tokens across all accessible web and print resources. Reaching standard pretraining targets (300 billion tokens) forces multiple training epochs ( extapprox8\ ext{approx } 8), causing parameter-heavy models (e.g. 13B parameters) to suffer performance degradation due to repeated-data overfitting.
    2. Securing high-quality prose required incorporating copyrighted materials from the National Library of Finland ( extapprox30,000\ ext{approx } 30,000 ePub eBooks and the Lehdet archived news collection, comprising 19.3 billion characters), which cannot be publicly redistributed. Consequently, full exact replication of the pretraining pipeline cannot be performed without independent bilateral access agreements with the national library.

Coverage note — No substantial contributed material was omitted; the knowls cover all model architectures, pretraining configurations, dataset curation and filtering pipelines, FIN-bench tasks, scaling results, BLUUMI transfer evaluation, toxicity, bias, alignment, and data scarcity limitations.

References

  1. 1.Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al. 2023. Santacoder: don’t reach for the stars! arXiv preprint arXiv:2301.03988.
  2. 2.Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021. A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861.
  3. 3.Giusepppe Attardi. 2015. Wikiextractor. https://github.com/attardi/wikiextractor.
  4. 4.Adrien Barbaresi. 2021. Trafilatura: A web scraping library and command-line tool for text discovery and extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 122–131.
  5. 5.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623.
  6. 6.Douglas Biber. 1988. Variation across speech and writing. Cambridge University Press, Cambridge.
  7. 7.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29.
  8. 8.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  10. 10.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  11. 11.Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. arXiv preprint arXiv:1805.10190.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  13. 13.Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Shrivastava, Samson Tan, et al. 2021. Nl-augmenter: A framework for task-sensitive natural language augmentation. arXiv preprint arXiv:2112.02721.
  14. 14.Anni Eskelinen, Laura Silvala, Filip Ginter, Sampo Pyysalo, and Veronika Laippala. 2023. Toxicity detection in Finnish using machine translation. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 685–697, Tórshavn, Faroe Islands. University of Tartu Library.
  15. 15.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. A framework for few-shot language model evaluation.
  16. 16.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  17. 17.Väinö Hatanpää. 2022. A generative pre-trained transformer model for Finnish. Master’s thesis, Aalto University. School of Science.
  18. 18.Kenneth Heafield. 2011. KenLM: Faster and smaller language model queries. In Proceedings of the sixth workshop on statistical machine translation, pages 187–197.
  19. 19.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  20. 20.Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas, Valtteri Skantsi, Jemina Kilpeläinen, Hanna-Mari Kupari, Jenna Saarni, Maija Sevón, and Otto Tarkka. 2021. Finnish paraphrase corpus. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa 2021).
  21. 21.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022. The BigScience ROOTS corpus: A 1.6 tb composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  22. 22.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. 2023. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161.
  23. 23.Juhani Luotolahti, Jenna Kanerva, Veronika Laippala, Sampo Pyysalo, and Filip Ginter. 2015. Towards universal web parsebanks. In International Conference on Dependency Linguistics.
  24. 24.Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems, volume 26. Curran Associates, Inc.
  25. 25.Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. 2023a. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124.
  26. 26.Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023b. Scaling data-constrained language models. arXiv preprint arXiv:2305.16264.
  27. 27.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  28. 28.Emily Öhman, Marc Pàmies, Kaisla Kajava, and Jörg Tiedemann. 2020. XED: A multilingual dataset for sentiment analysis and emotion detection. arXiv preprint arXiv:2011.01612.
  29. 29.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.
  30. 30.Jan Pomikálek. 2011. Removing boilerplate and duplicate content from web corpora. Ph.D. thesis, Masaryk university, Faculty of informatics, Brno, Czech Republic.
  31. 31.Ofir Press, Noah A Smith, and Mike Lewis. 2021. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409.
  32. 32.Sampo Pyysalo, Jenna Kanerva, Antti Virtanen, and Filip Ginter. 2021. Wikibert models: Deep transfer learning for many languages. In Proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa), pages 1–10.
  33. 33.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training.
  34. 34.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  35. 35.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506.
  36. 36.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022a. Bloom: A 176b-parameter open-access multilingual language model.
  37. 37.Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. 2022b. What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424.
  38. 38.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053.
  39. 39.Valtteri Skantsi and Veronika Laippala. 2022. Analyzing the unrestricted web: The finnish corpus of online registers. Nordic Journal of Linguistics.
  40. 40.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  41. 41.Erich Strohmaier, Jack Dongarra, Horst Simon, Martin Meuer, and Hans Meuer. 2023. Top500 - the list. https://www.top500.org/.
  42. 42.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  43. 43.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems, 30.
  44. 44.Viljami Venekoski and Jouko Vankka. 2017. Finnish resources for evaluating language model semantics. In Proceedings of the 21st Nordic Conference on Computational Linguistics, NoDaLiDa, 22-24 May 2017, Gothenburg, Sweden, 131, pages 231–236. Linköping University Electronic Press, Linköpings universitet.
  45. 45.Antti Virtanen, Jenna Kanerva, Rami Ilo, Jouni Luoma, Juhani Luotolahti, Tapio Salakoski, Filip Ginter, and Sampo Pyysalo. 2019. Multilingual is not enough: Bert for finnish. arXiv preprint arXiv:1912.07076.
  46. 46.Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359.
  47. 47.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498.
  48. 48.Zheng-Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, et al. 2022. BLOOM+ 1: Adding language support to bloom for zero-shot prompting. arXiv preprint arXiv:2212.09535.

Citation

MLA
Luukkonen, R., et al. “FinGPT: Large Generative Models for a Small Language”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2710–26, https://doi.org/10.18653/v1/2023.emnlp-main.164.
APA
Luukkonen, R., Komulainen, V., Luoma, J., Eskelinen, A., Kanerva, J., Kupari, H.-M., Ginter, F., Laippala, V., Muennighoff, N., Piktus, A., Wang, T., Tazi, N., Scao, T., Wolf, T., Suominen, O., Sairanen, S., Merioksa, M., Heinonen, J., Vahtola, A., … Pyysalo, S. (2023). FinGPT: Large Generative Models for a Small Language. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2710–2726. https://doi.org/10.18653/v1/2023.emnlp-main.164
Chicago
Luukkonen, R., V. Komulainen, J. Luoma, et al. 2023. “FinGPT: Large Generative Models for a Small Language”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2710–26. https://doi.org/10.18653/v1/2023.emnlp-main.164.
Harvard
Luukkonen, R. et al. (2023) “FinGPT: Large Generative Models for a Small Language”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 2710–2726. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.164.
Vancouver
1. Luukkonen R, Komulainen V, Luoma J, et al (2023) FinGPT: Large Generative Models for a Small Language. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 2710–2726

BibTeX

@inproceedings{luukkonen-etal-2023-fingpt,
    title = "{F}in{GPT}: Large Generative Models for a Small Language",
    author = "Luukkonen, Risto  and
      Komulainen, Ville  and
      Luoma, Jouni  and
      Eskelinen, Anni  and
      Kanerva, Jenna  and
      Kupari, Hanna-Mari  and
      Ginter, Filip  and
      Laippala, Veronika  and
      Muennighoff, Niklas  and
      Piktus, Aleksandra  and
      Wang, Thomas  and
      Tazi, Nouamane  and
      Scao, Teven  and
      Wolf, Thomas  and
      Suominen, Osma  and
      Sairanen, Samuli  and
      Merioksa, Mikko  and
      Heinonen, Jyrki  and
      Vahtola, Aija  and
      Antao, Samuel  and
      Pyysalo, Sampo",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.164/",
    doi = "10.18653/v1/2023.emnlp-main.164",
    pages = "2710--2726"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/