Predicate-Argument Structure Divergences in Chinese and English Parallel Sentences and their Impact on Language Transfer
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ASFDI97HN" target="_blank" >RIV/00216208:11320/26:SFDI97HN - isvavai.cz</a>
Result on the web
<a href="http://arxiv.org/abs/2511.09796" target="_blank" >http://arxiv.org/abs/2511.09796</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.48550/arXiv.2511.09796" target="_blank" >10.48550/arXiv.2511.09796</a>
Alternative languages
Result language
angličtina
Original language name
Predicate-Argument Structure Divergences in Chinese and English Parallel Sentences and their Impact on Language Transfer
Original language description
Cross-lingual Natural Language Processing (NLP) has gained significant traction in recent years, offering practical solutions in low-resource settings by transferring linguistic knowledge from resource-rich to low-resource languages. This field leverages techniques like annotation projection and model transfer for language adaptation, supported by multilingual pre-trained language models. However, linguistic divergences hinder language transfer, especially among typologically distant languages. In this paper, we present an analysis of predicate-argument structures in parallel Chinese and English sentences. We explore the alignment and misalignment of predicate annotations, inspecting similarities and differences and proposing a categorization of structural divergences. The analysis and the categorization are supported by a qualitative and quantitative analysis of the results of an annotation projection experiment, in which, in turn, one of the two languages has been used as source language to project annotations into the corresponding parallel sentences. The results of this analysis show clearly that language transfer is asymmetric. An aspect that requires attention when it comes to selecting the source language in transfer learning applications and that needs to be investigated before any scientific claim about cross-lingual NLP is proposed.
Czech name
—
Czech description
—
Classification
Type
O - Miscellaneous
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů