Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ASRCF7U3B" target="_blank" >RIV/00216208:11320/26:SRCF7U3B - isvavai.cz</a>
Result on the web
<a href="https://journals.openedition.org/corpus/9177" target="_blank" >https://journals.openedition.org/corpus/9177</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.4000/13654" target="_blank" >10.4000/13654</a>
Alternative languages
Result language
angličtina
Original language name
Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian
Original language description
Although Dagur and Alsatian represent two typologically distant language families, they share several similarities: both languages are endangered, do not have a unified spelling system, and have few available digital corpora. Given these challenges, the main aim of this article is to compare the noise in corpora for these languages and its impact on part-of-speech (POS) annotation and tagging. We first discuss what strategies can be used to reduce the noise due to spelling inconsistencies observed during corpus collection, using Dagur as an example. We then observe that the distributions of POS trigrams in the manually annotated Dagur and Alsatian corpora are similar to those of typologically related languages in UD v2.12, laying the foundations for experimenting with zero-shot transfer approaches for automatic POS tagging. We evaluate some simple noise reduction strategies for POS tagging using the example of the Alsatian dialects and relying on their proximity to standard German. The results obtained confirm the important role of linguistic proximity in automatic POS tagging and the effectiveness of the proposed data transformation method. However, they also invite further interpretation of the capacities of multilingual models.
Czech name
—
Czech description
—
Classification
Type
J<sub>ost</sub> - Miscellaneous article in a specialist periodical
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Corpus
ISSN
1765-3126
e-ISSN
—
Volume of the periodical
2025
Issue of the periodical within the volume
26
Country of publishing house
US - UNITED STATES
Number of pages
19
Pages from-to
1-19
UT code for WoS article
—
EID of the result in the Scopus database
—