Developing new annotated corpora for Lithuanian: Compilation issues
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A8RHGCNFW" target="_blank" >RIV/00216208:11320/26:8RHGCNFW - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >http://dx.doi.org/10.5755/j01.sal.1.46.40544</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.5755/j01.sal.1.46.40544" target="_blank" >10.5755/j01.sal.1.46.40544</a>
Alternative languages
Result language
angličtina
Original language name
Developing new annotated corpora for Lithuanian: Compilation issues
Original language description
The currently available Lithuanian grammatically annotated corpora (the morphologically annotated corpus MATAS and the syntactically annotated corpus ALKSNIS) are insufficient to meet the growing demands of Lithuanian language processing. Therefore, within the framework of the European Union’s NextGenerationEU project, “Morphologically and syntactically annotated text models for training (gold standards)”, two new corpora are being developed in accordance with the international Universal Dependencies (UD) standard. These new corpora (10 million tokens each) will represent the written variant of the Lithuanian language and, similar to the Corpus of Contemporary Lithuanian Language and other annotated corpora, they will consist of four sub-corpora containing texts from fiction, non-fiction (scientific literature), administrative texts, and online journalistic articles. Following the presentation of the annotated corpora MATAS and ALKSNIS, as well as the UD corpora of other languages, this paper discusses various aspects of the annotated corpora under development. First, the corpora structure, balance and sampling are described. Second, tokenization, essential for the automatic corpus analysis is examined, and third, the concept of token is discussed. Although corpora size can be expressed in words, it is most commonly measured in tokens (suggested Lithuanian translation: tekstyno vienetas), as corpora include not only words, but also non-words, e.g., punctuation marks, numerals, abbreviations, symbols, etc. Since the elements included under the concept of token may vary depending on the language and researchers’ decisions, the article discusses what is considered a token for Lithuanian within the boundaries of the current project. © 2025 Kauno Technologijos Universitetas. All rights reserved.
Czech name
—
Czech description
—
Classification
Type
J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Studies About Languages
ISSN
1648-2824
e-ISSN
—
Volume of the periodical
2025
Issue of the periodical within the volume
46
Country of publishing house
US - UNITED STATES
Number of pages
17
Pages from-to
119-135
UT code for WoS article
—
EID of the result in the Scopus database
2-s2.0-105012299913