Developing new annotated corpora for Lithuanian: Compilation issues
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A8RHGCNFW" target="_blank" >RIV/00216208:11320/26:8RHGCNFW - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >http://dx.doi.org/10.5755/j01.sal.1.46.40544</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >10.5755/j01.sal.1.46.40544</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Developing new annotated corpora for Lithuanian: Compilation issues
Popis výsledku v původním jazyce
The currently available Lithuanian grammatically annotated corpora (the morphologically annotated corpus MATAS and the syntactically annotated corpus ALKSNIS) are insufficient to meet the growing demands of Lithuanian language processing. Therefore, within the framework of the European Union’s NextGenerationEU project, “Morphologically and syntactically annotated text models for training (gold standards)”, two new corpora are being developed in accordance with the international Universal Dependencies (UD) standard. These new corpora (10 million tokens each) will represent the written variant of the Lithuanian language and, similar to the Corpus of Contemporary Lithuanian Language and other annotated corpora, they will consist of four sub-corpora containing texts from fiction, non-fiction (scientific literature), administrative texts, and online journalistic articles. Following the presentation of the annotated corpora MATAS and ALKSNIS, as well as the UD corpora of other languages, this paper discusses various aspects of the annotated corpora under development. First, the corpora structure, balance and sampling are described. Second, tokenization, essential for the automatic corpus analysis is examined, and third, the concept of token is discussed. Although corpora size can be expressed in words, it is most commonly measured in tokens (suggested Lithuanian translation: tekstyno vienetas), as corpora include not only words, but also non-words, e.g., punctuation marks, numerals, abbreviations, symbols, etc. Since the elements included under the concept of token may vary depending on the language and researchers’ decisions, the article discusses what is considered a token for Lithuanian within the boundaries of the current project. © 2025 Kauno Technologijos Universitetas. All rights reserved.
Název v anglickém jazyce
Developing new annotated corpora for Lithuanian: Compilation issues
Popis výsledku anglicky
The currently available Lithuanian grammatically annotated corpora (the morphologically annotated corpus MATAS and the syntactically annotated corpus ALKSNIS) are insufficient to meet the growing demands of Lithuanian language processing. Therefore, within the framework of the European Union’s NextGenerationEU project, “Morphologically and syntactically annotated text models for training (gold standards)”, two new corpora are being developed in accordance with the international Universal Dependencies (UD) standard. These new corpora (10 million tokens each) will represent the written variant of the Lithuanian language and, similar to the Corpus of Contemporary Lithuanian Language and other annotated corpora, they will consist of four sub-corpora containing texts from fiction, non-fiction (scientific literature), administrative texts, and online journalistic articles. Following the presentation of the annotated corpora MATAS and ALKSNIS, as well as the UD corpora of other languages, this paper discusses various aspects of the annotated corpora under development. First, the corpora structure, balance and sampling are described. Second, tokenization, essential for the automatic corpus analysis is examined, and third, the concept of token is discussed. Although corpora size can be expressed in words, it is most commonly measured in tokens (suggested Lithuanian translation: tekstyno vienetas), as corpora include not only words, but also non-words, e.g., punctuation marks, numerals, abbreviations, symbols, etc. Since the elements included under the concept of token may vary depending on the language and researchers’ decisions, the article discusses what is considered a token for Lithuanian within the boundaries of the current project. © 2025 Kauno Technologijos Universitetas. All rights reserved.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Studies About Languages
ISSN
1648-2824
e-ISSN
—
Svazek periodika
2025
Číslo periodika v rámci svazku
46
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
17
Strana od-do
119-135
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105012299913