An Attempt to Develop a Neural Parser Based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AFMFGBNBB" target="_blank" >RIV/00216208:11320/26:FMFGBNBB - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/978-981-96-4282-3_26" target="_blank" >http://dx.doi.org/10.1007/978-981-96-4282-3_26</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-981-96-4282-3_26" target="_blank" >10.1007/978-981-96-4282-3_26</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
An Attempt to Develop a Neural Parser Based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
Popis výsledku v původním jazyce
In this paper, we aimed to develop a neural parser for Vietnamese based on simplified Head-Driven Phrase Structure Grammar (HPSG). The existing corpora, VietTreebank and VnDT, had around 15% of constituency and dependency tree pairs that did not adhere to simplified HPSG rules. To attempt to address the issue of the corpora not adhering to simplified HPSG rules, we randomly permuted samples from the training and development sets to make them compliant with simplified HPSG. We then modified the first simplified HPSG Neural Parser for the Penn Treebank by replacing it with the PhoBERT or XLM-RoBERTa models, which can encode Vietnamese texts. We conducted experiments on our modified VietTreebank and VnDT corpora. Our extensive experiments showed that the simplified HPSG Neural Parser achieved a new state-of-the-art F-score of 82% for constituency parsing when using the same predicted part-of-speech (POS) tags as the self-attentive constituency parser. Additionally, it outperformed previous studies in dependency parsing with a higher Unlabeled Attachment Score (UAS). However, our parser obtained lower Labeled Attachment Score (LAS) scores likely due to our focus on arc permutation without changing the original labels, as we did not consult with a linguistic expert. Lastly, the research findings of this paper suggest that simplified HPSG should be given more attention to linguistic expert when developing treebanks for Vietnamese natural language processing. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
Název v anglickém jazyce
An Attempt to Develop a Neural Parser Based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
Popis výsledku anglicky
In this paper, we aimed to develop a neural parser for Vietnamese based on simplified Head-Driven Phrase Structure Grammar (HPSG). The existing corpora, VietTreebank and VnDT, had around 15% of constituency and dependency tree pairs that did not adhere to simplified HPSG rules. To attempt to address the issue of the corpora not adhering to simplified HPSG rules, we randomly permuted samples from the training and development sets to make them compliant with simplified HPSG. We then modified the first simplified HPSG Neural Parser for the Penn Treebank by replacing it with the PhoBERT or XLM-RoBERTa models, which can encode Vietnamese texts. We conducted experiments on our modified VietTreebank and VnDT corpora. Our extensive experiments showed that the simplified HPSG Neural Parser achieved a new state-of-the-art F-score of 82% for constituency parsing when using the same predicted part-of-speech (POS) tags as the self-attentive constituency parser. Additionally, it outperformed previous studies in dependency parsing with a higher Unlabeled Attachment Score (UAS). However, our parser obtained lower Labeled Attachment Score (LAS) scores likely due to our focus on arc permutation without changing the original labels, as we did not consult with a linguistic expert. Lastly, the research findings of this paper suggest that simplified HPSG should be given more attention to linguistic expert when developing treebanks for Vietnamese natural language processing. © The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Commun. Comput. Info. Sci.
ISBN
978-981-96-4281-6
ISSN
—
e-ISSN
—
Počet stran výsledku
16
Strana od-do
313-328
Název nakladatele
Springer Science and Business Media Deutschland GmbH
Místo vydání
—
Místo konání akce
Danang
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—