Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Domain Generalization in Vietnamese Dependency Parsing: A Novel Benchmark and Domain Gap Analysis

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AZ3IVR7F7" target="_blank" >RIV/00216208:11320/26:Z3IVR7F7 - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/978-981-96-4282-3_14" target="_blank" >http://dx.doi.org/10.1007/978-981-96-4282-3_14</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-981-96-4282-3_14" target="_blank" >10.1007/978-981-96-4282-3_14</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Domain Generalization in Vietnamese Dependency Parsing: A Novel Benchmark and Domain Gap Analysis

  • Popis výsledku v původním jazyce

    Dependency parsing has received significant attention from the research community due to its recognized applications across diverse areas of natural language processing (NLP). However, the majority of dependency parsing studies to date have not addressed the out-of-domain problem, where the data in the testing phase are in a different distribution compared with data in training domains, despite this being a common problem in practice. Furthermore, Vietnamese is still considered a low-resource language in parsing tasks, as most standard treebanks are primarily developed for more widely spoken languages such as English and Chinese. This shortage pushes the difficulty of studies of Vietnamese dependency parsing task even further. To advance research on domain generalization in Vietnamese dependency parsing task, this paper introduces a new treebank called DGDT(VietnameseDomainGeneralizationDependencyTreebank), where domains in train/dev/test set are completely separated. This is the distinction of our treebank, compared to other Vietnamese dependency treebanks. We also release DGDTMark, a cross-domain Vietnamese dependency parsing benchmark suite using our treebank to assess the generalization ability of parsers over domains. Moreover, our suite can support further research in analyzing the impacts of domain gaps on the dependency parsing task. Through experiments, we observe that the performance of parsers is most affected by two gaps: newspaper topics and writing styles. Besides, the performance drops remarkably by 3.27% UAS and 5.09% LAS in the scenario with the largest domain gap, which proves that our treebank poses a significant challenge for further research. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

  • Název v anglickém jazyce

    Domain Generalization in Vietnamese Dependency Parsing: A Novel Benchmark and Domain Gap Analysis

  • Popis výsledku anglicky

    Dependency parsing has received significant attention from the research community due to its recognized applications across diverse areas of natural language processing (NLP). However, the majority of dependency parsing studies to date have not addressed the out-of-domain problem, where the data in the testing phase are in a different distribution compared with data in training domains, despite this being a common problem in practice. Furthermore, Vietnamese is still considered a low-resource language in parsing tasks, as most standard treebanks are primarily developed for more widely spoken languages such as English and Chinese. This shortage pushes the difficulty of studies of Vietnamese dependency parsing task even further. To advance research on domain generalization in Vietnamese dependency parsing task, this paper introduces a new treebank called DGDT(VietnameseDomainGeneralizationDependencyTreebank), where domains in train/dev/test set are completely separated. This is the distinction of our treebank, compared to other Vietnamese dependency treebanks. We also release DGDTMark, a cross-domain Vietnamese dependency parsing benchmark suite using our treebank to assess the generalization ability of parsers over domains. Moreover, our suite can support further research in analyzing the impacts of domain gaps on the dependency parsing task. Through experiments, we observe that the performance of parsers is most affected by two gaps: newspaper topics and writing styles. Besides, the performance drops remarkably by 3.27% UAS and 5.09% LAS in the scenario with the largest domain gap, which proves that our treebank poses a significant challenge for further research. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

  • Projekt

  • Návaznosti

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    Commun. Comput. Info. Sci.

  • ISBN

    978-981-96-4281-6

  • ISSN

  • e-ISSN

  • Počet stran výsledku

    15

  • Strana od-do

    167-181

  • Název nakladatele

    Springer Science and Business Media Deutschland GmbH

  • Místo vydání

  • Místo konání akce

    Danang

  • Datum konání akce

    1. 1. 2026

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku