Distribution of dependency tags in different text types in Czech
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61988987%3A17250%2F25%3AA2603CPM" target="_blank" >RIV/61988987:17250/25:A2603CPM - isvavai.cz</a>
Alternative codes found
RIV/00216208:90244/25:10513827
Result on the web
<a href="https://benjamins.com/catalog/cilt.370.08che" target="_blank" >https://benjamins.com/catalog/cilt.370.08che</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1075/cilt.370.08che" target="_blank" >10.1075/cilt.370.08che</a>
Alternative languages
Result language
angličtina
Original language name
Distribution of dependency tags in different text types in Czech
Original language description
The objective of this chapter is to examine how dependency tags are distributed across various genres in the Czech language. Specifically, our focus is on identifying the similarities and differences in these distributions among different types of text. This study utilizes data from the SYN2020, a large balanced corpus of contemporary written Czech, comprising 100 million words. The findings indicate that the frequencies of dependency tags follow a power-law distribution, represented by the equation y = ax-b. When analyzing the values of the parameters a and b, a distinct pattern emerges that differentiates the genres. Genres within fiction, such as poetry, drama, novels, and short stories, typically exhibit lower values for a and b, whereas non-fiction genres, including administrative texts and professional literature, demonstrate higher values. Journalistic texts, like newspapers and leisure magazines, fall in between fiction and non-fiction literature in terms of these parameter values. Consequently, comparing these parameters appears to be an effective method for conducting stylometric research.
Czech name
—
Czech description
—
Classification
Type
C - Chapter in a specialist book
CEP classification
—
OECD FORD branch
60203 - Linguistics
Result continuities
Project
<a href="/en/project/GA22-20632S" target="_blank" >GA22-20632S: Quantitative Syntactic Stylistics of Contemporary Written Czech</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Book/collection name
Mathematical Modelling in Linguistics and Text Analysis
ISBN
9789027228376
Number of pages of the result
14
Pages from-to
90-103
Number of pages of the book
241
Publisher name
John Benjamins Publishing Company
Place of publication
Amsterdam
UT code for WoS chapter
001609792500009