Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Evaluation of the morphological rules for the Tenyidie language: a low-resource language

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ANL4QDGKA" target="_blank" >RIV/00216208:11320/26:NL4QDGKA - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/s10579-024-09788-y" target="_blank" >http://dx.doi.org/10.1007/s10579-024-09788-y</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/s10579-024-09788-y" target="_blank" >10.1007/s10579-024-09788-y</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Evaluation of the morphological rules for the Tenyidie language: a low-resource language

  • Popis výsledku v původním jazyce

    The Tenyidie language, a.k.a Angami language, is a low-resource language belonging to the Tibeto-Burman Language family, which is spoken by the Tenyimia Community and is considered a major language in Nagaland in the north-eastern part of India. Tenyidie is tonal, SOV, and highly agglutinative in its linguistics characteristics. Among the Natural Language Processing (NLP) tasks, part-of-speech (POS) tagging is one of the primary tasks that is used in building many other NLP tasks, such as dependency parsing, named entity recognition, machine translation, etc. The main aim of this paper is to evaluate the morphological rules in the Tenyidie Language by building a morphological rule-based POS tagger. In this work, the morphological rules have been evaluated on 158,403 annotated tokens in Tenyidie. To the best of the authors knowledge, there is no reported work on the evaluation of the morphological rules for the Tenyidie Language. The main contributions of this research are the evaluation of the existing morphological rules in the Tenyidie Language by building a morphological rule-based POS tagger, and the creation of 158,403 tokens annotated dataset for evaluating the morphological rules. In addition, we have introduced some new morphological rules. © The Author(s), under exclusive licence to Springer Nature B.V. 2024.

  • Název v anglickém jazyce

    Evaluation of the morphological rules for the Tenyidie language: a low-resource language

  • Popis výsledku anglicky

    The Tenyidie language, a.k.a Angami language, is a low-resource language belonging to the Tibeto-Burman Language family, which is spoken by the Tenyimia Community and is considered a major language in Nagaland in the north-eastern part of India. Tenyidie is tonal, SOV, and highly agglutinative in its linguistics characteristics. Among the Natural Language Processing (NLP) tasks, part-of-speech (POS) tagging is one of the primary tasks that is used in building many other NLP tasks, such as dependency parsing, named entity recognition, machine translation, etc. The main aim of this paper is to evaluate the morphological rules in the Tenyidie Language by building a morphological rule-based POS tagger. In this work, the morphological rules have been evaluated on 158,403 annotated tokens in Tenyidie. To the best of the authors knowledge, there is no reported work on the evaluation of the morphological rules for the Tenyidie Language. The main contributions of this research are the evaluation of the existing morphological rules in the Tenyidie Language by building a morphological rule-based POS tagger, and the creation of 158,403 tokens annotated dataset for evaluating the morphological rules. In addition, we have introduced some new morphological rules. © The Author(s), under exclusive licence to Springer Nature B.V. 2024.

Klasifikace

  • Druh

    J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Language Resources and Evaluation

  • ISSN

    1574-020X

  • e-ISSN

  • Svazek periodika

    59

  • Číslo periodika v rámci svazku

    3

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    26

  • Strana od-do

    3189-3214

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus

    2-s2.0-85210475503