All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Language Representation Models for Low- and Medium-Resource Languages

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AIUV97UZF" target="_blank" >RIV/00216208:11320/26:IUV97UZF - isvavai.cz</a>

  • Result on the web

    <a href="https://hdl.handle.net/20.500.11815/5957" target="_blank" >https://hdl.handle.net/20.500.11815/5957</a>

  • DOI - Digital Object Identifier

Alternative languages

  • Result language

    angličtina

  • Original language name

    Language Representation Models for Low- and Medium-Resource Languages

  • Original language description

    Transformer-based language models have proven to be extremely effective for a wide variety of natural language understanding tasks, including question answering, automatic text summarization, and sentiment analysis. These models are typically pre-trained on large, unannotated corpora using self-supervised tasks such as masked token prediction, often requiring weeks or months of training, followed by fine-tuning on practical tasks, which requires substantially less time and data by comparison. Since their introduction, Transformer models have grown exponentially in size, from approximately 100 million parameters in 2018 to over 600 billion in 2024, with the largest pre-training corpora growing from around 800 million tokens to over 14.8 trillion. However, many low- and medium-resource languages lack the extensive datasets and computational resources required to pre-train language models at this scale. Therefore, data-efficient pre-training techniques are crucial for effectively utilizing the limited resources available for these languages. In this thesis, we investigate various data-efficient pre-training strategies and evaluate their impact on downstream tasks in six low- to medium-resource languages: Icelandic, Estonian, Basque, Galician, Nepali, and Tajik. First, we analyze several text quality filtering techniques to discard noisy data from web-crawled corpora. We propose a novel, language-independent filtering approach using unsupervised clustering and outlier detection algorithms which achieves comparable performance to a rule-based approach. Second, we explore the effects of augmenting monolingual pre-training corpora with text from related and unrelated languages, as well as Python code, finding significant improvements in downstream performance for certain tasks for larger models. Our results support the hypothesis that linguistic similarity facilitates cross-lingual transfer. Finally, we compare several subword tokenization algorithms and evaluate their impact on downstream results when used in pre-trained language models. Our analysis reveals that the Unigram algorithm consistently yields the best results on downstream tasks, and that a vocabulary size of 64k outperforms smaller vocabularies by a statistically significant margin. Our findings demonstrate that data-efficient pre-training techniques can substantially improve the performance of language models for low- and medium-resource languages. By optimizing the use of available data and resources, we achieve statistically significant improvements in downstream tasks under data-constrained conditions, paving the way for more effective natural language processing in resource-constrained settings. We release several datasets and tools compiled and developed during the work of this thesis, as well as multiple pre-trained Transformer-based language models.

  • Czech name

  • Czech description

Classification

  • Type

    B - Specialist book

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • ISBN

    978-9935-539-78-6

  • Number of pages

    104

  • Publisher name

    Department of Computer Science, Reykjavík University

  • Place of publication

    Reykjavík, Island

  • UT code for WoS book