FinGPT: Large Generative Models for a Small Language
Risto LuukkonenVille KomulainenJouni LuomaAnni EskelinenJenna KanervaHanna-Mari KupariFilip GinterVeronika LaippalaNiklas MuennighoffAleksandra Piktus
Presents open-source Finnish generative language models scaling up to 176 billion parameters alongside a localized evaluation benchmark, offering a practical framework for developing large-scale foundation models for lower-resourced languages.
Most modern large language models require hundreds of billions of words of training text and disproportionately focus on English, leaving smaller languages underserved. Finnish is natively spoken by fewer than six million people and accounts for less than one percent of common online text repositories, making the development of capable generative artificial intelligence challenging under severe data constraints.
The article demonstrates effective strategies for developing advanced generative language models for data-constrained languages by training seven monolingual Finnish models from scratch and adapting an existing massive multilingual model for Finnish capabilities.
To establish credibility and broad linguistic coverage, the authors compiled a 207-billion-character dataset combining web crawls, national news archives, discussion forums, social media, and electronic books. After filtering for quality, perplexity, and toxicity, the preprocessed corpus yielded 38 billion tokens. Using the LUMI supercomputer, the authors trained seven monolingual models ranging from 186 million to 13 billion parameters over roughly eight passes (300 billion tokens). Additionally, they performed continued training on the 176-billion-parameter multilingual BLOOM model using a mix of its original training data and Finnish text, resulting in an extended model named BLUUMI. Model evaluation was conducted using FIN-bench, a newly adapted benchmark comprising 3,919 Finnish test examples across multiple reasoning and language tasks.
The analysis produced several key findings. First, the 176-billion-parameter BLUUMI model achieved the highest overall benchmark performance, outperforming the best prior Finnish model by over 20 percentage points and the original multilingual model by 12 to 18 percentage points on Finnish tasks, without degrading its English proficiency. Second, the monolingual 8-billion-parameter model delivered the strongest performance among the models trained from scratch, outperforming previous Finnish models of comparable size by over 10 percentage points. Third, performance declined between the 8-billion and 13-billion-parameter monolingual models, indicating that roughly 10 billion parameters represents an upper limit when training from scratch on limited text before overfitting reduces capability. Fourth, toxicity filtering of training data reduced the rate of toxic text generation by more than half compared to unfiltered models, although models still generated toxic output roughly 2% of the time. Finally, all evaluated base models performed poorly on basic alignment tests measuring helpfulness, honesty, and harmlessness.
These findings prove that high-performing foundation models can be successfully developed for lesser-resourced languages through rigorous data curation and extended pretraining of existing multilingual models. The results indicate that adapting existing large multilingual models is more effective than training large monolingual models from scratch when language data is limited. However, because these systems retain occupational gender biases and fail baseline safety and helpfulness evaluations, deploying them in unaligned forms presents significant compliance, reputational, and operational risks.
Organizations should treat these released models strictly as research foundations rather than production-ready systems. They should not be deployed in user-facing settings or high-stakes decision workflows, such as hiring, without substantial secondary intervention. Stakeholders pursuing language model implementations in data-constrained languages should prioritize continued pretraining of existing multilingual models over building large standalone models from scratch, and should incorporate explicit instruction tuning and safety alignment before operational release.
The findings are subject to specific limitations. The available volume of unique Finnish text forced repetitive training over multiple epochs, which contributed to performance degradation in the largest monolingual model. Furthermore, full replication of the training corpus is restricted because high-quality electronic book and newspaper data from the National Library of Finland cannot be redistributed due to copyright. Consequently, readers can have high confidence in the benchmarked technical capabilities but must remain cautious regarding the models' readiness for direct deployment.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). Introduces the 176-billion-parameter open multilingual foundation model that FinGPT directly adapts via continued pretraining to construct BLUUMI.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Establishes compute-optimal parameter-to-token scaling ratios that provide the theoretical foundation for FinGPT's analysis of model capacity limits under data scarcity.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Provides the methodology and framing for evaluating toxic degeneration in language models that FinGPT applies to assess its data filtering pipeline.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). Generalizes continued pretraining techniques from a single low-resource language to scale multilingual models across more than 500 diverse languages.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). Extends the evaluation of generative language models to a comprehensive 70-language benchmark across diverse NLP tasks and prompting strategies.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). Introduces an internal-representation metric to systematically quantify and compare foundation model proficiency between high- and low-resource languages.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). Builds a standardized testbed to systematically isolate and measure how specific data filtering and curation pipelines impact downstream language model capability.
