Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ASRCF7U3B" target="_blank" >RIV/00216208:11320/26:SRCF7U3B - isvavai.cz</a>

  • Výsledek na webu

    <a href="https://journals.openedition.org/corpus/9177" target="_blank" >https://journals.openedition.org/corpus/9177</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.4000/13654" target="_blank" >10.4000/13654</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian

  • Popis výsledku v původním jazyce

    Although Dagur and Alsatian represent two typologically distant language families, they share several similarities: both languages are endangered, do not have a unified spelling system, and have few available digital corpora. Given these challenges, the main aim of this article is to compare the noise in corpora for these languages and its impact on part-of-speech (POS) annotation and tagging. We first discuss what strategies can be used to reduce the noise due to spelling inconsistencies observed during corpus collection, using Dagur as an example. We then observe that the distributions of POS trigrams in the manually annotated Dagur and Alsatian corpora are similar to those of typologically related languages in UD v2.12, laying the foundations for experimenting with zero-shot transfer approaches for automatic POS tagging. We evaluate some simple noise reduction strategies for POS tagging using the example of the Alsatian dialects and relying on their proximity to standard German. The results obtained confirm the important role of linguistic proximity in automatic POS tagging and the effectiveness of the proposed data transformation method. However, they also invite further interpretation of the capacities of multilingual models.

  • Název v anglickém jazyce

    Managing Noise in Part-of-Speech Tagging for Extremely Low-Resource Languages: Comparing Strategies for Corpus Collection and Annotation in Dagur and Alsatian

  • Popis výsledku anglicky

    Although Dagur and Alsatian represent two typologically distant language families, they share several similarities: both languages are endangered, do not have a unified spelling system, and have few available digital corpora. Given these challenges, the main aim of this article is to compare the noise in corpora for these languages and its impact on part-of-speech (POS) annotation and tagging. We first discuss what strategies can be used to reduce the noise due to spelling inconsistencies observed during corpus collection, using Dagur as an example. We then observe that the distributions of POS trigrams in the manually annotated Dagur and Alsatian corpora are similar to those of typologically related languages in UD v2.12, laying the foundations for experimenting with zero-shot transfer approaches for automatic POS tagging. We evaluate some simple noise reduction strategies for POS tagging using the example of the Alsatian dialects and relying on their proximity to standard German. The results obtained confirm the important role of linguistic proximity in automatic POS tagging and the effectiveness of the proposed data transformation method. However, they also invite further interpretation of the capacities of multilingual models.

Klasifikace

  • Druh

    J<sub>ost</sub> - Ostatní články v recenzovaných periodicích

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Corpus

  • ISSN

    1765-3126

  • e-ISSN

  • Svazek periodika

    2025

  • Číslo periodika v rámci svazku

    26

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    19

  • Strana od-do

    1-19

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus