Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Developing new annotated corpora for Lithuanian: Compilation issues

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A8RHGCNFW" target="_blank" >RIV/00216208:11320/26:8RHGCNFW - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >http://dx.doi.org/10.5755/j01.sal.1.46.40544</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >10.5755/j01.sal.1.46.40544</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Developing new annotated corpora for Lithuanian: Compilation issues

  • Popis výsledku v původním jazyce

    The currently available Lithuanian grammatically annotated corpora (the morphologically annotated corpus MATAS and the syntactically annotated corpus ALKSNIS) are insufficient to meet the growing demands of Lithuanian language processing. Therefore, within the framework of the European Union’s NextGenerationEU project, “Morphologically and syntactically annotated text models for training (gold standards)”, two new corpora are being developed in accordance with the international Universal Dependencies (UD) standard. These new corpora (10 million tokens each) will represent the written variant of the Lithuanian language and, similar to the Corpus of Contemporary Lithuanian Language and other annotated corpora, they will consist of four sub-corpora containing texts from fiction, non-fiction (scientific literature), administrative texts, and online journalistic articles. Following the presentation of the annotated corpora MATAS and ALKSNIS, as well as the UD corpora of other languages, this paper discusses various aspects of the annotated corpora under development. First, the corpora structure, balance and sampling are described. Second, tokenization, essential for the automatic corpus analysis is examined, and third, the concept of token is discussed. Although corpora size can be expressed in words, it is most commonly measured in tokens (suggested Lithuanian translation: tekstyno vienetas), as corpora include not only words, but also non-words, e.g., punctuation marks, numerals, abbreviations, symbols, etc. Since the elements included under the concept of token may vary depending on the language and researchers’ decisions, the article discusses what is considered a token for Lithuanian within the boundaries of the current project. © 2025 Kauno Technologijos Universitetas. All rights reserved.

  • Název v anglickém jazyce

    Developing new annotated corpora for Lithuanian: Compilation issues

  • Popis výsledku anglicky

    The currently available Lithuanian grammatically annotated corpora (the morphologically annotated corpus MATAS and the syntactically annotated corpus ALKSNIS) are insufficient to meet the growing demands of Lithuanian language processing. Therefore, within the framework of the European Union’s NextGenerationEU project, “Morphologically and syntactically annotated text models for training (gold standards)”, two new corpora are being developed in accordance with the international Universal Dependencies (UD) standard. These new corpora (10 million tokens each) will represent the written variant of the Lithuanian language and, similar to the Corpus of Contemporary Lithuanian Language and other annotated corpora, they will consist of four sub-corpora containing texts from fiction, non-fiction (scientific literature), administrative texts, and online journalistic articles. Following the presentation of the annotated corpora MATAS and ALKSNIS, as well as the UD corpora of other languages, this paper discusses various aspects of the annotated corpora under development. First, the corpora structure, balance and sampling are described. Second, tokenization, essential for the automatic corpus analysis is examined, and third, the concept of token is discussed. Although corpora size can be expressed in words, it is most commonly measured in tokens (suggested Lithuanian translation: tekstyno vienetas), as corpora include not only words, but also non-words, e.g., punctuation marks, numerals, abbreviations, symbols, etc. Since the elements included under the concept of token may vary depending on the language and researchers’ decisions, the article discusses what is considered a token for Lithuanian within the boundaries of the current project. © 2025 Kauno Technologijos Universitetas. All rights reserved.

Klasifikace

  • Druh

    J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Studies About Languages

  • ISSN

    1648-2824

  • e-ISSN

  • Svazek periodika

    2025

  • Číslo periodika v rámci svazku

    46

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    17

  • Strana od-do

    119-135

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus

    2-s2.0-105012299913