Measuring and Evaluating Syntactic Distance Across Languages Using Universal Dependencies
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AESWSKN4D" target="_blank" >RIV/00216208:11320/26:ESWSKN4D - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1080/09296174.2025.2569581" target="_blank" >http://dx.doi.org/10.1080/09296174.2025.2569581</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1080/09296174.2025.2569581" target="_blank" >10.1080/09296174.2025.2569581</a>
Alternative languages
Result language
angličtina
Original language name
Measuring and Evaluating Syntactic Distance Across Languages Using Universal Dependencies
Original language description
This study addresses the calculation and evaluation of syntactic distance, which is a quantitative measure of structural similarity or divergence between languages. Building on existing alignment-based, feature-based and data-driven approaches, we introduce a novel hypergraph-based metric that assesses syntactic distance through structural alignment while explicitly incorporating word order features. The approach is then applied to a multilingual parallel corpus annotated within the Universal Dependencies (UD) framework, yielding syntactic distances between English and 19 non-English languages. Empirical evaluation further demonstrates the robustness and effectiveness of the proposed measure. Compared with approaches that ablate the hypergraph formalism, ignore word order or rely solely on data-driven metrics, the new metric proves robust under random sampling variation and effectively captures syntactic distance: statistical analyses show that intra-group language pairs exhibit significantly shorter syntactic distances than inter-group pairs. This approach thus provides a novel, formally grounded perspective on language distance based purely on structural properties. © 2025 Informa UK Limited, trading as Taylor & Francis Group.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Journal of Quantitative Linguistics
ISSN
0929-6174
e-ISSN
—
Volume of the periodical
2025
Issue of the periodical within the volume
2025
Country of publishing house
US - UNITED STATES
Number of pages
22
Pages from-to
1-22
UT code for WoS article
001597624200001
EID of the result in the Scopus database
2-s2.0-105019711693