Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

High-Quality LLM Pre-Training Texts from Dictionary Data.

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00142889" target="_blank" >RIV/00216224:14330/25:00142889 - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://nlp.fi.muni.cz/raslan/raslan25.pdf" target="_blank" >https://nlp.fi.muni.cz/raslan/raslan25.pdf</a>

  • DOI - Digital Object Identifier

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    High-Quality LLM Pre-Training Texts from Dictionary Data.

  • Popis výsledku v původním jazyce

    The quality of the pre-training texts is an important aspect in the development of a Large Language Model (LLM). High-quality data, such as collections of textbooks, academic papers, and educational forums, has been shown to improve model performance, generalization, and reduce biases. However, obtaining such data at scale can be challenging, especially for non-mainstream languages like Czech. In this paper, we introduce a method for generating high-quality Czech pre-training data from structured dictionary resources. By employing retrieval-augmented prompting and open-source LLMs, we transform XML-encoded lexicographic dictionary entries into fluent, semantically rich text. The resulting dataset demonstrates that dictionary-grounded generation can effectively enhance data quality. We present the results of experiments with several LLMs and the process of creating a new Czech pre-training dataset, SlamaHQTrain. This dataset was obtained by processing eight Czech dictionaries containing more than 500,000 entries and 18 million words.

  • Název v anglickém jazyce

    High-Quality LLM Pre-Training Texts from Dictionary Data.

  • Popis výsledku anglicky

    The quality of the pre-training texts is an important aspect in the development of a Large Language Model (LLM). High-quality data, such as collections of textbooks, academic papers, and educational forums, has been shown to improve model performance, generalization, and reduce biases. However, obtaining such data at scale can be challenging, especially for non-mainstream languages like Czech. In this paper, we introduce a method for generating high-quality Czech pre-training data from structured dictionary resources. By employing retrieval-augmented prompting and open-source LLMs, we transform XML-encoded lexicographic dictionary entries into fluent, semantically rich text. The resulting dataset demonstrates that dictionary-grounded generation can effectively enhance data quality. We present the results of experiments with several LLMs and the process of creating a new Czech pre-training dataset, SlamaHQTrain. This dataset was obtained by processing eight Czech dictionaries containing more than 500,000 entries and 18 million words.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10200 - Computer and information sciences

Návaznosti výsledku

  • Projekt

    <a href="/cs/project/EH23_025%2F0008710" target="_blank" >EH23_025/0008710: Na všechno sami: příležitosti a rizika individualizace společnosti</a><br>

  • Návaznosti

    P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Recent Advances in Slavonic Natural Language Processing, RASLAN 2025

  • ISBN

    9788026318583

  • ISSN

    2336-4289

  • e-ISSN

  • Počet stran výsledku

    16

  • Strana od-do

    69-84

  • Název nakladatele

    Tribun EU

  • Místo vydání

    Brno, Czech Republic

  • Místo konání akce

    Kouty nad Desnou, Česká Republika

  • Datum konání akce

    1. 1. 2025

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku