The Thai Universal Dependency Treebank
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AV7K7V5H2" target="_blank" >RIV/00216208:11320/26:V7K7V5H2 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1162/tacl_a_00745" target="_blank" >http://dx.doi.org/10.1162/tacl_a_00745</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1162/tacl_a_00745" target="_blank" >10.1162/tacl_a_00745</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
The Thai Universal Dependency Treebank
Popis výsledku v původním jazyce
Automatic dependency parsing of Thai sentences has been underexplored, as evidenced by the lack of large Thai dependency treebanks with complete dependency structures and the lack of a published evaluation of state-of-the-art models, especially transformer-based parsers. In this work, we addressed these gaps by introducing the Thai Universal Dependency Treebank (TUD), a new Thai treebank consisting of 3,627 trees annotated according to the Universal Dependencies (UD) framework. We then benchmarked 92 dependency parsing models that incorporate pretrained transformers on Thai-PUD and our TUD, achieving state-of-the-art results and shedding light on the optimal model components for Thai dependency parsing. Our error analysis of the models also reveals that polyfunctional words, serial verb construction, and lack of rich morphosyntactic features present main challenges for Thai dependency parsing. © 2025 Association for Computational Linguistics.
Název v anglickém jazyce
The Thai Universal Dependency Treebank
Popis výsledku anglicky
Automatic dependency parsing of Thai sentences has been underexplored, as evidenced by the lack of large Thai dependency treebanks with complete dependency structures and the lack of a published evaluation of state-of-the-art models, especially transformer-based parsers. In this work, we addressed these gaps by introducing the Thai Universal Dependency Treebank (TUD), a new Thai treebank consisting of 3,627 trees annotated according to the Universal Dependencies (UD) framework. We then benchmarked 92 dependency parsing models that incorporate pretrained transformers on Thai-PUD and our TUD, achieving state-of-the-art results and shedding light on the optimal model components for Thai dependency parsing. Our error analysis of the models also reveals that polyfunctional words, serial verb construction, and lack of rich morphosyntactic features present main challenges for Thai dependency parsing. © 2025 Association for Computational Linguistics.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Transactions of the Association for Computational Linguistics
ISSN
2307-387X
e-ISSN
—
Svazek periodika
13
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
16
Strana od-do
376-391
Kód UT WoS článku
001511539600005
EID výsledku v databázi Scopus
2-s2.0-105010272565