Decomposing dependency analysis: revisiting the relation between annotation scheme and structure-based textual measures
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ARRQ4WCKM" target="_blank" >RIV/00216208:11320/26:RRQ4WCKM - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1093/llc/fqaf003" target="_blank" >http://dx.doi.org/10.1093/llc/fqaf003</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/llc/fqaf003" target="_blank" >10.1093/llc/fqaf003</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Decomposing dependency analysis: revisiting the relation between annotation scheme and structure-based textual measures
Popis výsledku v původním jazyce
Standardized quantitative measurement of texts lies at the heart of digital approaches to humanities. Structure-based textual measures are known to be influenced by the choice of syntactic annotation schemes. Building on previous research, the present article further explores the relation between annotation schemes and the index of mean dependency distance (MDD) by comparing the treebanks of seventeen languages, respectively, within a tree representation (basic universal dependencies, BUD) and within a graphic representation (enhanced universal dependencies, EUD). Following the idea of decomposing annotation schemes into the combinations of analyses of specific constructions (coordinate structures, control constructions, and relative clauses), we design algorithms to identify them in the CoNLL-U format treebanks and explore their influences. It is found that the overall MDD of the EUD representation is statistically higher than that of BUD at corpus level, primarily affected by the coordinate structure due to its high frequency. At sentence level, all three constructions might contribute to either increased or decreased MDD, with stochastically intervening words and word order being two important determinants of the values of the measure. Finally, we propose and argue for the view that MDDs calculated under different annotation schemes should be regarded as different textual measures in nature. In sum, the present study provides another case study to deepen our understanding of the nature of syntactic annotation schemes and its relation with textual indices, which paves the way for standard measurement of texts in future humanities research. © The Author(s) 2025. Published by Oxford University Press on behalf of EADH. All rights reserved.
Název v anglickém jazyce
Decomposing dependency analysis: revisiting the relation between annotation scheme and structure-based textual measures
Popis výsledku anglicky
Standardized quantitative measurement of texts lies at the heart of digital approaches to humanities. Structure-based textual measures are known to be influenced by the choice of syntactic annotation schemes. Building on previous research, the present article further explores the relation between annotation schemes and the index of mean dependency distance (MDD) by comparing the treebanks of seventeen languages, respectively, within a tree representation (basic universal dependencies, BUD) and within a graphic representation (enhanced universal dependencies, EUD). Following the idea of decomposing annotation schemes into the combinations of analyses of specific constructions (coordinate structures, control constructions, and relative clauses), we design algorithms to identify them in the CoNLL-U format treebanks and explore their influences. It is found that the overall MDD of the EUD representation is statistically higher than that of BUD at corpus level, primarily affected by the coordinate structure due to its high frequency. At sentence level, all three constructions might contribute to either increased or decreased MDD, with stochastically intervening words and word order being two important determinants of the values of the measure. Finally, we propose and argue for the view that MDDs calculated under different annotation schemes should be regarded as different textual measures in nature. In sum, the present study provides another case study to deepen our understanding of the nature of syntactic annotation schemes and its relation with textual indices, which paves the way for standard measurement of texts in future humanities research. © The Author(s) 2025. Published by Oxford University Press on behalf of EADH. All rights reserved.
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Digital Scholarship in the Humanities
ISSN
2055-7671
e-ISSN
—
Svazek periodika
40
Číslo periodika v rámci svazku
1
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
19
Strana od-do
400-418
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105001965945