Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

ACCURACY EVALUATION AND ERROR ANALYSIS OF DEPENDENCY PARSING FOR TEXTS IN UKRAINIAN

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A5VZ9MNLD" target="_blank" >RIV/00216208:11320/26:5VZ9MNLD - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.30837/2522-9818.2025.2.102" target="_blank" >http://dx.doi.org/10.30837/2522-9818.2025.2.102</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.30837/2522-9818.2025.2.102" target="_blank" >10.30837/2522-9818.2025.2.102</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    ACCURACY EVALUATION AND ERROR ANALYSIS OF DEPENDENCY PARSING FOR TEXTS IN UKRAINIAN

  • Popis výsledku v původním jazyce

    The subject of our research is the dependency parsing of sentences in the Ukrainian language using the Universal Dependencies framework. The goal of the work is to evaluate the accuracy of existing transition-based and graph-based parsing architectures with and without deep word embeddings on the Ukrainian dataset, and to analyze the error profiles of such parsers. The article addresses two tasks. One is to evaluate the accuracy of several modern dependency parsing approaches applied to a hand-annotated gold standard dataset, using labeled and unlabeled attachment scores as the metric to evaluate the parsing accuracy. The other task is to analyze and categorize the errors made by standard parsers. Resolving these errors could potentially allow us to build a more accurate parser in the future. Error rate for different categories is compared to the ba seline error rate, and statistical significance of such comparison is validated using the chi-square method. The key results are as follows. For the Ukrainian language, parsing accuracy is greatly increased with the use of deep word embeddings. Transition-based parser with deep word embeddings provides the highest labeled attachment score of 84.66% for the test dataset. For the same parser, higher error rates are associated with non-projectivity of dependencies, higher sentence length and higher distance to head. Also, for pronouns and numerals the error rate for labeled attachment is significantly higher than the baseline, while the unlabeled error rate is at the baseline. Conclusions: parsing accuracy for the Ukrainian dataset is sub-par in comparison with other languages, but the overall trend of accuracy improvement with the use of deep word embeddings is consistent with existing research. To improve overall parsing accuracy, we must focus on such problem areas as non-projective dependencies, longer sentences, and greater distance between the head and the dependent. In future work we intend to explore ways to improve parsing accuracy by supplementing neural parsing with other approaches, like formal rules or pre-and post-processing. © K. Syrotkin, 2025.

  • Název v anglickém jazyce

    ACCURACY EVALUATION AND ERROR ANALYSIS OF DEPENDENCY PARSING FOR TEXTS IN UKRAINIAN

  • Popis výsledku anglicky

    The subject of our research is the dependency parsing of sentences in the Ukrainian language using the Universal Dependencies framework. The goal of the work is to evaluate the accuracy of existing transition-based and graph-based parsing architectures with and without deep word embeddings on the Ukrainian dataset, and to analyze the error profiles of such parsers. The article addresses two tasks. One is to evaluate the accuracy of several modern dependency parsing approaches applied to a hand-annotated gold standard dataset, using labeled and unlabeled attachment scores as the metric to evaluate the parsing accuracy. The other task is to analyze and categorize the errors made by standard parsers. Resolving these errors could potentially allow us to build a more accurate parser in the future. Error rate for different categories is compared to the ba seline error rate, and statistical significance of such comparison is validated using the chi-square method. The key results are as follows. For the Ukrainian language, parsing accuracy is greatly increased with the use of deep word embeddings. Transition-based parser with deep word embeddings provides the highest labeled attachment score of 84.66% for the test dataset. For the same parser, higher error rates are associated with non-projectivity of dependencies, higher sentence length and higher distance to head. Also, for pronouns and numerals the error rate for labeled attachment is significantly higher than the baseline, while the unlabeled error rate is at the baseline. Conclusions: parsing accuracy for the Ukrainian dataset is sub-par in comparison with other languages, but the overall trend of accuracy improvement with the use of deep word embeddings is consistent with existing research. To improve overall parsing accuracy, we must focus on such problem areas as non-projective dependencies, longer sentences, and greater distance between the head and the dependent. In future work we intend to explore ways to improve parsing accuracy by supplementing neural parsing with other approaches, like formal rules or pre-and post-processing. © K. Syrotkin, 2025.

Klasifikace

  • Druh

    J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název periodika

    Innovative Technologies and Scientific Solutions for Industries

  • ISSN

    2522-9818

  • e-ISSN

  • Svazek periodika

    2025

  • Číslo periodika v rámci svazku

    2

  • Stát vydavatele periodika

    US - Spojené státy americké

  • Počet stran výsledku

    9

  • Strana od-do

    102-110

  • Kód UT WoS článku

  • EID výsledku v databázi Scopus

    2-s2.0-105019070600