Distribution of dependency tags in different text types in Czech
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61988987%3A17250%2F25%3AA2603CPM" target="_blank" >RIV/61988987:17250/25:A2603CPM - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:90244/25:10513827
Výsledek na webu
<a href="https://benjamins.com/catalog/cilt.370.08che" target="_blank" >https://benjamins.com/catalog/cilt.370.08che</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1075/cilt.370.08che" target="_blank" >10.1075/cilt.370.08che</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Distribution of dependency tags in different text types in Czech
Popis výsledku v původním jazyce
The objective of this chapter is to examine how dependency tags are distributed across various genres in the Czech language. Specifically, our focus is on identifying the similarities and differences in these distributions among different types of text. This study utilizes data from the SYN2020, a large balanced corpus of contemporary written Czech, comprising 100 million words. The findings indicate that the frequencies of dependency tags follow a power-law distribution, represented by the equation y = ax-b. When analyzing the values of the parameters a and b, a distinct pattern emerges that differentiates the genres. Genres within fiction, such as poetry, drama, novels, and short stories, typically exhibit lower values for a and b, whereas non-fiction genres, including administrative texts and professional literature, demonstrate higher values. Journalistic texts, like newspapers and leisure magazines, fall in between fiction and non-fiction literature in terms of these parameter values. Consequently, comparing these parameters appears to be an effective method for conducting stylometric research.
Název v anglickém jazyce
Distribution of dependency tags in different text types in Czech
Popis výsledku anglicky
The objective of this chapter is to examine how dependency tags are distributed across various genres in the Czech language. Specifically, our focus is on identifying the similarities and differences in these distributions among different types of text. This study utilizes data from the SYN2020, a large balanced corpus of contemporary written Czech, comprising 100 million words. The findings indicate that the frequencies of dependency tags follow a power-law distribution, represented by the equation y = ax-b. When analyzing the values of the parameters a and b, a distinct pattern emerges that differentiates the genres. Genres within fiction, such as poetry, drama, novels, and short stories, typically exhibit lower values for a and b, whereas non-fiction genres, including administrative texts and professional literature, demonstrate higher values. Journalistic texts, like newspapers and leisure magazines, fall in between fiction and non-fiction literature in terms of these parameter values. Consequently, comparing these parameters appears to be an effective method for conducting stylometric research.
Klasifikace
Druh
C - Kapitola v odborné knize
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
<a href="/cs/project/GA22-20632S" target="_blank" >GA22-20632S: Kvantitativní syntaktická stylistika současné psané češtiny</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název knihy nebo sborníku
Mathematical Modelling in Linguistics and Text Analysis
ISBN
9789027228376
Počet stran výsledku
14
Strana od-do
90-103
Počet stran knihy
241
Název nakladatele
John Benjamins Publishing Company
Místo vydání
Amsterdam
Kód UT WoS kapitoly
001609792500009