Syntactic units and their length distributions: A case study in Czech
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61988987%3A17250%2F25%3AA2603D5A" target="_blank" >RIV/61988987:17250/25:A2603D5A - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:11320/26:QSFFW4ZB RIV/00216208:90244/25:10513792
Výsledek na webu
<a href="https://aclanthology.org/volumes/2025.quasy-1/" target="_blank" >https://aclanthology.org/volumes/2025.quasy-1/</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Syntactic units and their length distributions: A case study in Czech
Popis výsledku v původním jazyce
This study investigates the length distributions of syntactic units in Czech across multiple hierarchical levels: sentences, independent clauses, clauses, phrases, subphrases, and chunks. Using a diverse dataset – including Universal Dependency treebanks, presidential speeches, the Czech Bible, and random sample from corpora of modern Czech – the analysis examines whether lengths of these syntactic units follow consistent distributional patterns. Length is defined as the number of immediate subunits, and the distributions were modeled using the hyper-Poisson distribution. The results demonstrate that the hyper-Poisson model fits well distributions of length of all abovementioned syntactic units, pointing to a common principle underlying the organization of syntactic structure in Czech.
Název v anglickém jazyce
Syntactic units and their length distributions: A case study in Czech
Popis výsledku anglicky
This study investigates the length distributions of syntactic units in Czech across multiple hierarchical levels: sentences, independent clauses, clauses, phrases, subphrases, and chunks. Using a diverse dataset – including Universal Dependency treebanks, presidential speeches, the Czech Bible, and random sample from corpora of modern Czech – the analysis examines whether lengths of these syntactic units follow consistent distributional patterns. Length is defined as the number of immediate subunits, and the distributions were modeled using the hyper-Poisson distribution. The results demonstrate that the hyper-Poisson model fits well distributions of length of all abovementioned syntactic units, pointing to a common principle underlying the organization of syntactic structure in Czech.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
60203 - Linguistics
Návaznosti výsledku
Projekt
—
Návaznosti
V - Vyzkumna aktivita podporovana z jinych verejnych zdroju
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the Third Workshop on Quantitative Syntax (QUASY, SyntaxFest 2025)
ISBN
979-8-89176-293-0
ISSN
—
e-ISSN
—
Počet stran výsledku
9
Strana od-do
115-123
Název nakladatele
Association for Computational Linguistics
Místo vydání
Kerrville
Místo konání akce
Ljubljana
Datum konání akce
29. 8. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—