All

What are you looking for?

All
Projects
Results
Organizations

Quick search

  • Projects supported by TA ČR
  • Excellent projects
  • Projects with the highest public support
  • Current projects

Smart search

  • That is how I find a specific +word
  • That is how I leave the -word out of the results
  • “That is how I can find the whole phrase”

Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian

The result's identifiers

  • Result code in IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ASRCF7U3B" target="_blank" >RIV/00216208:11320/26:SRCF7U3B - isvavai.cz</a>

  • Result on the web

    <a href="https://journals.openedition.org/corpus/9177" target="_blank" >https://journals.openedition.org/corpus/9177</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.4000/13654" target="_blank" >10.4000/13654</a>

Alternative languages

  • Result language

    angličtina

  • Original language name

    Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian

  • Original language description

    Although Dagur and Alsatian represent two typologically distant language families, they share several similarities: both languages are endangered, do not have a unified spelling system, and have few available digital corpora. Given these challenges, the main aim of this article is to compare the noise in corpora for these languages and its impact on part-of-speech (POS) annotation and tagging. We first discuss what strategies can be used to reduce the noise due to spelling inconsistencies observed during corpus collection, using Dagur as an example. We then observe that the distributions of POS trigrams in the manually annotated Dagur and Alsatian corpora are similar to those of typologically related languages in UD v2.12, laying the foundations for experimenting with zero-shot transfer approaches for automatic POS tagging. We evaluate some simple noise reduction strategies for POS tagging using the example of the Alsatian dialects and relying on their proximity to standard German. The results obtained confirm the important role of linguistic proximity in automatic POS tagging and the effectiveness of the proposed data transformation method. However, they also invite further interpretation of the capacities of multilingual models.

  • Czech name

  • Czech description

Classification

  • Type

    J<sub>ost</sub> - Miscellaneous article in a specialist periodical

  • CEP classification

  • OECD FORD branch

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

  • Project

  • Continuities

Others

  • Publication year

    2025

  • Confidentiality

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

  • Name of the periodical

    Corpus

  • ISSN

    1765-3126

  • e-ISSN

  • Volume of the periodical

    2025

  • Issue of the periodical within the volume

    26

  • Country of publishing house

    US - UNITED STATES

  • Number of pages

    19

  • Pages from-to

    1-19

  • UT code for WoS article

  • EID of the result in the Scopus database