Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3A9J9UNDVG" target="_blank" >RIV/00216208:11320/26:9J9UNDVG - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1145/3748315" target="_blank" >http://dx.doi.org/10.1145/3748315</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1145/3748315" target="_blank" >10.1145/3748315</a>
Alternative languages
Result language
angličtina
Original language name
Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing
Original language description
Dependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions. © 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Czech name
—
Czech description
—
Classification
Type
J<sub>SC</sub> - Article in a specialist periodical, which is included in the SCOPUS database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
ACM Transactions on Asian and Low-Resource Language Information Processing
ISSN
2375-4699
e-ISSN
—
Volume of the periodical
24
Issue of the periodical within the volume
8
Country of publishing house
US - UNITED STATES
Number of pages
9
Pages from-to
83
UT code for WoS article
—
EID of the result in the Scopus database
2-s2.0-105018459708