Measuring and Evaluating Syntactic Distance Across Languages Using Universal Dependencies
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AESWSKN4D" target="_blank" >RIV/00216208:11320/26:ESWSKN4D - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1080/09296174.2025.2569581" target="_blank" >http://dx.doi.org/10.1080/09296174.2025.2569581</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1080/09296174.2025.2569581" target="_blank" >10.1080/09296174.2025.2569581</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Measuring and Evaluating Syntactic Distance Across Languages Using Universal Dependencies
Popis výsledku v původním jazyce
This study addresses the calculation and evaluation of syntactic distance, which is a quantitative measure of structural similarity or divergence between languages. Building on existing alignment-based, feature-based and data-driven approaches, we introduce a novel hypergraph-based metric that assesses syntactic distance through structural alignment while explicitly incorporating word order features. The approach is then applied to a multilingual parallel corpus annotated within the Universal Dependencies (UD) framework, yielding syntactic distances between English and 19 non-English languages. Empirical evaluation further demonstrates the robustness and effectiveness of the proposed measure. Compared with approaches that ablate the hypergraph formalism, ignore word order or rely solely on data-driven metrics, the new metric proves robust under random sampling variation and effectively captures syntactic distance: statistical analyses show that intra-group language pairs exhibit significantly shorter syntactic distances than inter-group pairs. This approach thus provides a novel, formally grounded perspective on language distance based purely on structural properties. © 2025 Informa UK Limited, trading as Taylor & Francis Group.
Název v anglickém jazyce
Measuring and Evaluating Syntactic Distance Across Languages Using Universal Dependencies
Popis výsledku anglicky
This study addresses the calculation and evaluation of syntactic distance, which is a quantitative measure of structural similarity or divergence between languages. Building on existing alignment-based, feature-based and data-driven approaches, we introduce a novel hypergraph-based metric that assesses syntactic distance through structural alignment while explicitly incorporating word order features. The approach is then applied to a multilingual parallel corpus annotated within the Universal Dependencies (UD) framework, yielding syntactic distances between English and 19 non-English languages. Empirical evaluation further demonstrates the robustness and effectiveness of the proposed measure. Compared with approaches that ablate the hypergraph formalism, ignore word order or rely solely on data-driven metrics, the new metric proves robust under random sampling variation and effectively captures syntactic distance: statistical analyses show that intra-group language pairs exhibit significantly shorter syntactic distances than inter-group pairs. This approach thus provides a novel, formally grounded perspective on language distance based purely on structural properties. © 2025 Informa UK Limited, trading as Taylor & Francis Group.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Journal of Quantitative Linguistics
ISSN
0929-6174
e-ISSN
—
Svazek periodika
2025
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
22
Strana od-do
1-22
Kód UT WoS článku
001597624200001
EID výsledku v databázi Scopus
2-s2.0-105019711693