Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A9J9UNDVG" target="_blank" >RIV/00216208:11320/26:9J9UNDVG - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1145/3748315" target="_blank" >http://dx.doi.org/10.1145/3748315</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1145/3748315" target="_blank" >10.1145/3748315</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
Popis výsledku v původním jazyce
Dependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions. © 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Název v anglickém jazyce
Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
Popis výsledku anglicky
Dependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions. © 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
ACM Transactions on Asian and Low-Resource Language Information Processing
ISSN
2375-4699
e-ISSN
—
Svazek periodika
24
Číslo periodika v rámci svazku
8
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
9
Strana od-do
83
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105018459708